Open the selector to load the current language catalog.
Translations are provided by Google Translate.
Cost simulator
Subscription versus API, priced from sources
Compare a source-priced subscription against a modelled API workload. A price override is always labelled as your assumption, and any applicable rate the source did not publish makes the estimate unavailable instead of cheap.
1
Subscription plan
Price, billing cadence, minimum seats, and availability all come from the plans source.
Up to four models sharing one workload, so the shares always total 100%. Changing one moves the others in proportion; with a single model it is the whole workload.
Token volumes are your assumptions, not source evidence. The long-context buffer scales both input and output tokens. Enter the workload for the whole team: seats multiply the subscription price only, and never the API token volume.
Conversations / dayHow many separate chats or agent runs the whole team starts on a working day. Multiplied by messages per conversation and active days to get monthly messages.
Messages / conversationTurns in one conversation, counting both sides. A long thread costs more than a short one because each turn resends the history.
Active days / monthWorking days per month this workload actually runs. 22 is a weekday month; use 30 for something that runs every day.
Input tokens / messageAverage prompt size for one turn, including any system prompt and resent history. Priced at the input rate, and the number the long-context buffer scales.
Output tokens / messageAverage completion size for one turn. Priced at the output rate, which is typically several times the input rate — this is usually the input worth checking first.
Cache read share (%)Percentage of input tokens served from a prompt cache, priced at the cache read rate. The remainder is billed at the full input rate.
Cache write share (%)Percentage of input tokens written into a prompt cache, priced at the cache write rate. Read and write together cannot exceed 100%; if a selected model publishes no cache rate, the estimate reports that rather than pricing it at zero.
Long-context buffer (%)Headroom for long-context turns, as a percentage. Scales both input and output tokens per message before anything is priced — 50% means every turn is modelled at one and a half times the sizes above.
Character → token estimator
Planning assumption: 4 characters ≈ 1 token for English. Ratios differ by roughly threefold across languages, so a single figure would be misleading — but none of them replaces the provider’s own tokenizer, which is the only authority on what you will be billed for.
4 chars / token≈ 0 tokens
Subscription
$20.00
/ month · source plan price
No token allowance equivalence is assumed for a subscription.
API estimate
$56.32
/ month · 2.82M modelled tokens
Source rates × the workload and mix you entered.
Breakeven crossover
1M
modelled tokens / month
Subscription cost ÷ this scenario’s effective API rate ($20.00 / 1M).
0–1B token cost curve
The API line uses this scenario’s effective blended mix rate per million tokens. The subscription line is flat because no token allowance equivalence is claimed for a plan.
Itemised source prices
Exactly what each source published. Nothing on this table is derived.