ling

InclusionAI Ling 3.0 Flash

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.

Fresh · fetched 2026-08-17T02:11:57.739Z · Models.dev(opens in a new tab)

Minimum context256K
Output256K
Providers5
CapabilitiesTool calling, Reasoning

Provider offers

ProvidersInput / 1MOutput / 1MCache read / 1MContext / outputUpdated
NanoGPTinclusionai/ling-3.0-flash$0.075$0.22$0.015256K2026-07-23
OpenRouterinclusionai/ling-3.0-flash$0.021$0.063$0.004256K2026-07-23
Vercel AI Gatewayinclusionai/ling-3.0-flash$0.060$0.18$0.012256K2026-08-06
LLM Gatewayling-3.0-flash$0.060$0.18$0.012256K2026-08-02
Kilo Gatewayinclusionai/ling-3.0-flash$0.060$0.18$0.012256K2026-07-23

Lowest reported provider price in USD per 1M tokens; unavailable is not free. Provider documentation remains authoritative. Catalogue presence is not an uptime, security, or quality endorsement.

Source metadata and limitations

Capabilities are source-reported. Tool calling does not prove quality in a particular harness. Prices and limits can change.

Modalities: texttext