inclusionAI shipped Ling-3.0-flash, a fast execution model for the grunt nodes in an agent graph

If you run agents, the notable drop this week is an execution model rather than another would-be planner: Ant's inclusionAI shipped Ling-3.0-flash (it lists as AntLing-3.0-flash on OpenRouter). It's built for the Loop/Router/single-step-executor nodes where you want low latency and tool calls that don't fall apart, not deep reasoning. Specs that matter for that job: 124B total but only 5.1B active (sparse MoE, so it's cheap and quick), 256K context, TTFT under 100ms, and a thinking mode you can toggle with enable_thinking. The part I'd actually care about for agents is the long-horizon tool calling — they RL'd it hard on tool-call checking, so it holds up across a multi-step run instead of drifting halfway through.


This is a companion discussion topic for the original entry at https://www.reddit.com/r/LocalLLM/comments/1v740a9/inclusionai_shipped_ling30flash_a_fast_execution/