Unknown
Llama 3.2 11b Vision Instruct
Open Llama multimodal model for image understanding and text reasoning
Fresh · fetched 2026-08-17T02:11:57.739Z · Models.dev(opens in a new tab)
Minimum context128K
Output4K
Providers2
CapabilitiesTool calling, Image input
Provider offers
| Providers | Input / 1M | Output / 1M | Cache read / 1M | Context / output | Updated |
|---|---|---|---|---|---|
meta/llama-3.2-11b-vision-instruct | $0.00 | $0.00 | Unavailable | 128K | 2024-09-18 |
meta/llama-3.2-11b-vision-instruct | $0.055 | $0.055 | Unavailable | 16K | 2025-01-01 |
Lowest reported provider price in USD per 1M tokens; unavailable is not free. Provider documentation remains authoritative. Catalogue presence is not an uptime, security, or quality endorsement.
Source metadata and limitations
Capabilities are source-reported. Tool calling does not prove quality in a particular harness. Prices and limits can change.
Modalities: text, image → text