Ant Group open-sources Ling-3.0-flash-VL: native multimodal model, scores 25 on Artificial Analysis Intelligence Index
Ant Group's first native multimodal model in the Ling line, Ling-3.0-flash-VL, has been open-sourced on Hugging Face and ModelScope with 124B total / 5.5B active params, 256K context, and BF16+FP8 weights.
Ant Group has open-sourced Ling-3.0-flash-VL, the first native multimodal model in its Ling line, via its inclusionAI organization on Hugging Face and ModelScope, with BF16 and FP8 weights both made public. Built on the Ling-3.0-flash mixture-of-experts backbone, the model has 124B total parameters and activates roughly 5.5B per token. It natively accepts image, text, and video input with a 256K-token context window. On the Artificial Analysis Intelligence Index v4.3, Ling-3.0-flash-VL scored 25, meaningfully above the 16 posted by Qwen3.5 122B A10B at a comparable active-parameter size. Ant's own model card reports a 42 on the v4.1.1 protocol; the gap reflects different evaluation sets rather than a contradiction. The Ling line is the model family that Ant's foundational intelligence unit has been opening up since August 2026, covering Ling-3.0-flash, Ling-3.0-tiny, the finance-tuned Ling-3.0-flash-Fin, and now the first native multimodal release. Ant published an SGLang deployment cookbook alongside the weights and maintains a Ling-tuned vLLM fork for the architecture. Ling-3.0-flash-Fin, released earlier, is the first finance-augmented model in the Ling line. It keeps the 124B-total/5.1B-active configuration and a 256K context window, with continued pretraining on financial corpora and tool-use optimization for annual reports, financial workbooks, and multi-document research workflows.