ling
InclusionAI Ling 3.0 Flash
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.
Fresh · fetched 2026-08-17T02:11:57.739Z · Models.dev(opens in a new tab)
Minimum context256K
Output256K
Providers5
CapabilitiesTool calling, Reasoning
Provider offers
| Providers | Input / 1M | Output / 1M | Cache read / 1M | Context / output | Updated |
|---|---|---|---|---|---|
inclusionai/ling-3.0-flash | $0.075 | $0.22 | $0.015 | 256K | 2026-07-23 |
inclusionai/ling-3.0-flash | $0.021 | $0.063 | $0.004 | 256K | 2026-07-23 |
inclusionai/ling-3.0-flash | $0.060 | $0.18 | $0.012 | 256K | 2026-08-06 |
ling-3.0-flash | $0.060 | $0.18 | $0.012 | 256K | 2026-08-02 |
inclusionai/ling-3.0-flash | $0.060 | $0.18 | $0.012 | 256K | 2026-07-23 |
Lowest reported provider price in USD per 1M tokens; unavailable is not free. Provider documentation remains authoritative. Catalogue presence is not an uptime, security, or quality endorsement.
Source metadata and limitations
Capabilities are source-reported. Tool calling does not prove quality in a particular harness. Prices and limits can change.
Modalities: text → text