Tencent Hunyuan Vision API: Multimodal at ¥2.85 (~$0.40) per M Tokens (2026)
Two Hunyuan vision tiers: T1 reasons before answering, 1.5 Instruct answers straight. Both ¥2.85/¥8.55 per M tokens — among the cheapest vision-capable models anywhere. Best-in-class for Chinese OCR.
Picking between the two
One question: does the answer require thinking? Chart-trend interpretation, object-relationship inference, look-and-judge tasks — T1 (analyzes first, steadier, slightly slower). Text extraction, recognition, template tagging — 1.5 Instruct (direct, fast). Same price, so switch freely per task.
The sweet spot
E-commerce image tagging, receipt/screenshot OCR, content-moderation prescreening, video-frame understanding — many images, simple per-call asks, huge volume. At ¥2.85 input, a million tokens of image descriptions costs about forty cents; running such loads on flagship multimodals is pure waste.
The 24K boundary, honestly
24K-28K does not fit long mixed documents. The right shape is a one-image-one-question pipeline, not dropping in a PDF (PDF input is unsupported on our relay anyway — docs section 8). For long-context multimodal, use Gemini 3.1 Pro or Qwen3.5-VL.
Setup
OpenAI-compatible: base_url https://cocodot.co/api/ai/v1, model hunyuan-t1-vision-20250916 or hunyuan-vision-1.5-instruct, images in standard OpenAI vision format. The $0.5 signup credit covers it — test on your own images before topping up.
Hunyuan vision tiers on cocodot (95% of list, CNY per M tokens)
| Model | In / Out | Context | Best for |
|---|---|---|---|
| Hunyuan T1 Vision | ¥2.85 / ¥8.55 | 28K | reasoned visual QA, charts |
| Hunyuan Vision 1.5 | ¥2.85 / ¥8.55 | 24K | OCR, recognition, batch tagging |