"""MiniMax H3 text/vision conditioning: Qwen3-VL-32B (truncated to 50 layers). The H3 presentation is NOT chat-templated: token ids are raw prompt/label text (no special tokens) with explicit vision blocks spliced in: t2va: fl2va: ": " [": " ] ref2va: per condition in request order (1-based ordinals per type): image -> ": " audio -> "