Unknown

pixtral-12b-2409

Pixtral 2409 12B is a state-of-the-art multimodal model with 12B parameters and a 400M vision encoder, natively trained on interleaved text and image data. It excels in tasks spanning vision-language reasoning, instruction following, and pure text understanding, making it highly effective for real-world multimodal applications.

Fresh · fetched 2026-08-17T02:11:57.739Z · Models.dev(opens in a new tab)

Minimum context128K
Output128K
Providers4
CapabilitiesTool calling, Reasoning, Image input

Provider offers

ProvidersInput / 1MOutput / 1MCache read / 1MContext / outputUpdated
Cortecspixtral-12b-2409$0.22$0.22Unavailable128K2024-11-09
GreenPTpixtral-12b-2409$0.28$0.28Unavailable128K2024-09-01
Scalewaypixtral-12b-2409$0.20$0.20Unavailable128K2026-03-17
Pioneermistralai/Pixtral-12B-2409$0.15$0.15$0.15128K2024-09-01

Lowest reported provider price in USD per 1M tokens; unavailable is not free. Provider documentation remains authoritative. Catalogue presence is not an uptime, security, or quality endorsement.

Source metadata and limitations

Capabilities are source-reported. Tool calling does not prove quality in a particular harness. Prices and limits can change.

Modalities: text, imagetext