Skip to main content
ModelScale

Cost simulator

Subscription versus API, priced from sources

Compare a source-priced subscription against a modelled API workload. A price override is always labelled as your assumption, and any applicable rate the source did not publish makes the estimate unavailable instead of cheap.

Subscription plan

Price, billing cadence, minimum seats, and availability all come from the plans source.

$20.00 monthly chargeMinimum seats: none publishedAvailability: availableAPI billed separatelyOpenAI Help Center · verified 2026-09-17

API model mix

Up to four models sharing one workload, so the shares always total 100%. Changing one moves the others in proportion; with a single model it is the whole workload.

Monthly workload

Token volumes are your assumptions, not source evidence. The long-context buffer scales both input and output tokens. Enter the workload for the whole team: seats multiply the subscription price only, and never the API token volume.

Conversations / dayHow many separate chats or agent runs the whole team starts on a working day. Multiplied by messages per conversation and active days to get monthly messages.
Messages / conversationTurns in one conversation, counting both sides. A long thread costs more than a short one because each turn resends the history.
Active days / monthWorking days per month this workload actually runs. 22 is a weekday month; use 30 for something that runs every day.
Input tokens / messageAverage prompt size for one turn, including any system prompt and resent history. Priced at the input rate, and the number the long-context buffer scales.
Output tokens / messageAverage completion size for one turn. Priced at the output rate, which is typically several times the input rate — this is usually the input worth checking first.
Cache read share (%)Percentage of input tokens served from a prompt cache, priced at the cache read rate. The remainder is billed at the full input rate.
Cache write share (%)Percentage of input tokens written into a prompt cache, priced at the cache write rate. Read and write together cannot exceed 100%; if a selected model publishes no cache rate, the estimate reports that rather than pricing it at zero.
Long-context buffer (%)Headroom for long-context turns, as a percentage. Scales both input and output tokens per message before anything is priced — 50% means every turn is modelled at one and a half times the sizes above.
Character → token estimator

Planning assumption: 4 characters ≈ 1 token for English. Ratios differ by roughly threefold across languages, so a single figure would be misleading — but none of them replaces the provider’s own tokenizer, which is the only authority on what you will be billed for.

4 chars / token0 tokens

Subscription

$20.00

/ month · source plan price

No token allowance equivalence is assumed for a subscription.

API estimate

$56.32

/ month · 2.82M modelled tokens

Source rates × the workload and mix you entered.

Breakeven crossover

1M

modelled tokens / month

Subscription cost ÷ this scenario’s effective API rate ($20.00 / 1M).

0–1B token cost curve

The API line uses this scenario’s effective blended mix rate per million tokens. The subscription line is flat because no token allowance equivalence is claimed for a plan.

Itemised source prices

Exactly what each source published. Nothing on this table is derived.

Source prices for the selected plan and models
ItemComponentSource valueEvidence
OpenAI ChatGPT Plusmonthly charge$20.00OpenAI Help Center · verified 2026-09-17
effective monthly / seat$20.00
Claude Fable 5.1Input / 1M$10.00
Output / 1M$50.00
Cache read / 1M$0.25
Cache write / 1MUnavailable

Derived monthly line items

Everything here is computed from your workload and the source prices above.

Derived monthly cost per route
RouteShareModelled tokensMonthly costStatus
Claude Fable 5.1100%2.82M$56.32Priced from source rates
API total100%2.82M$56.32Sum of the routes above
Subscription1 seats$20.00Source plan price × seats