StackPricing
All AI models

Mistral · AI model API pricing

How much does Mistral Small 4 cost?

A serious budget contender from Europe — between DeepSeek V4 Flash and Gemini Flash-Lite on price.

Rate card

Input tokens$0.15 / 1M
Cached input$0.015 / 1M
Output tokens$0.60 / 1M
Context window

What that means per month

WorkloadTokens / monthCost / month
Prototype 1M in · 0.3M out $0.33/mo
Small app 10M in · 3M out $3.30/mo
Production app 50M in · 15M out $16.5/mo
High volume 200M in · 60M out $66/mo
Heavy platform 1B in · 300M out $330/mo
Price it at your exact usage Official pricing page

Standard on-demand rates from Mistral's official pricing page, last checked 2026-07-27. Cached input reads bill at 10% of the input rate, with no cache-creation or storage fee. Monthly figures exclude caching and batch discounts. Confirm live pricing before committing.

Where Mistral Small 4 sits on price

At the reference month of 10M input and 3M output tokens, Mistral Small 4 costs $3.30, which makes it the 3rd cheapest of the 30 models tracked here. That is 1.9× the price of Qwen-Flash, the cheapest model at this mix ($1.70). Among the 9 budget-tier models it ranks 3rd, behind Qwen-Flash at $1.70.

Cache economics

Cached input on Mistral Small 4 bills at $0.015 per 1M — 10% of its own input rate, against a 10% median across the models here that publish a cached-read price. 2 models discount cached reads more steeply than this. On the reference month, serving 80% of that input from cache takes the input side of the bill from $1.50 to $0.42, leaving output untouched at $1.80 — which is why a high reuse rate changes the ranking above rather than just shaving the total.

Cache reads only. Writes are billed separately by some providers and are excluded here — see the rate note above for Mistral's terms. Try it at your own hit rate on the AI cost calculator.

Cheaper alternatives

At a reference workload of 10M input / 3M output tokens a month, these cost less than Mistral Small 4 ($3.30/mo):

Mistral Small 4 head-to-head