AI Is Sold by the Token. Here Is What That Means for Your Bill.
AI APIs charge per token, and output costs three to six times more than input. Prices fell roughly 80% between early 2025 and early 2026, from under $0.20 per million tokens to $10 at the flagship end.
By the UISC BD Editorial Desk · United Information Service Center · Published 13 September 2026 · 6-minute read
Anyone building with AI eventually meets a bill priced in "tokens per million". Few people explain what a token is, or why the same task can cost a hundred times more on one model than another.
What a Token Is
Models do not read words or letters. They read tokens — chunks of text that are often a word, part of a word, or a punctuation mark.
In English, a rough rule of thumb is that a token is about three-quarters of a word. Languages written in scripts that are less represented in training data, including Bangla, typically break into more tokens for the same meaning — which means the same message can cost more to process.
Why Output Costs More Than Input
Nearly every provider charges 3 to 6 times more for output tokens than input tokens.
The reason is computational. Reading your prompt can be processed largely in parallel. Generating an answer happens one token at a time, each step depending on the last. Writing is the expensive part.
The practical lesson: a long document with a short answer is cheap. A short question with a long answer is not.
What It Costs in 2026
Prices per million tokens, input / output, as listed in September 2026 price comparisons:
- Qwen3.7 Flash: $0.03 / $0.13 — the cheapest listed paid API.
- Gemini 2.5 Flash-Lite: $0.10 / $0.40.
- DeepSeek-V4: $0.14 / $0.28.
- Mistral Small 4: $0.15 / $0.60.
- GPT-5.6 Luna: $0.20 / $1.20, after a 30 July price cut.
- Mid-tier models cluster around $2 input — Claude Sonnet 5 was listed at introductory pricing of $2 / $10 through 31 August.
- Flagships: GPT-5.6 Sol $5 / $30, Claude Opus 5 $5 / $25, Claude Fable 5 $10 / $50.
Prices change often. Treat these as a snapshot of the shape of the market, and check a provider's own page before budgeting.
The 80 Percent Collapse
LLM API prices fell approximately 80 percent between early 2025 and early 2026. That is one of the steepest price declines for any widely used technology service.
It happened because of competition and efficiency together. As covered in our report on four frontier launches in one week, new models are now shipping at the same price as the ones they replace, and competition between providers has tightened.
Where the Money Actually Goes
Long context. Filling a 1-million-token context costs about $0.14 on DeepSeek V4 Flash and about $10 on Claude Fable 5 — a 71-times spread. See our context window explainer for why stuffing everything in is rarely the right move anyway.
Agents. An AI agent may make dozens of model calls to finish one task. Cheap per-call prices multiply quickly.
Conversation history. In most chat applications, the whole conversation is re-sent with every new message. A long thread gets more expensive with each reply.
Five Ways to Cut the Bill
- Match the model to the task. Classification and extraction rarely need a flagship.
- Ask for shorter answers. Output is the expensive side.
- Retrieve, do not stuff. Send the relevant passages, not the whole archive — the principle behind RAG.
- Start fresh conversations. Stop paying to resend old history.
- Consider local models. For routine work, a model on your own machine costs electricity, not tokens.
Related reading
- Why AI Forgets: Context Windows Explained Plainly
- Four Frontier AI Models Launched in 72 Hours This Month
- How to Run AI on Your Own Computer: Ollama and llama.cpp
- How AI Image Generators Work, and How to Get the Picture You Want
Sources
- "LLM API pricing comparison and calculator (September 2026)," BenchLM.ai — benchlm.ai
- "LLM API pricing comparison in 2026: every major model ranked by cost," CloudZero — cloudzero.com
- "LLM API pricing 2026: GPT vs Claude vs Gemini vs DeepSeek," Spheron — spheron.network
- "AI model context window comparison 2026: advertised vs. real," elvex — elvex.com