gemma

Gemma 4 12B Instruct

Google's Gemma 4 12B Instruct is an open-weight multimodal model for text, image, audio, and video understanding, with tool calling and structured output support.

Fresh · fetched 2026-08-17T02:11:57.739Z · Models.dev(opens in a new tab)

Minimum context256K
Output32K
Providers2
CapabilitiesTool calling, Reasoning, Image input

Provider offers

ProvidersInput / 1MOutput / 1MCache read / 1MContext / outputUpdated
NanoGPTgemma-4-12b-it$0.060$0.30$0.030256K2026-08-01
Pioneergoogle/gemma-4-12B-it$0.25$0.25$0.2532K2026-05-31

Lowest reported provider price in USD per 1M tokens; unavailable is not free. Provider documentation remains authoritative. Catalogue presence is not an uptime, security, or quality endorsement.

Source metadata and limitations

Capabilities are source-reported. Tool calling does not prove quality in a particular harness. Prices and limits can change.

Modalities: text, image, video, audiotext