AI model tracker, September 2026
Eight labs, twenty-two current models, one table. Every row was checked against the maker's own documentation or pricing page in September 2026; where a figure could not be verified we say so or left the row out. Prices are per million tokens on the standard API tier, US dollars unless marked.

- Read time9 min
- Data tables2
- StatusUpdated September 2026
Figures come from vendor pricing pages, model cards and release notes read in September 2026, cross-checked against independent price trackers. Some vendor links on this site are affiliate links; they never affect what we list. See our disclosure.
The table
| Model | Maker | Released | Context | API in / out per 1M | Free for consumers | Best for |
|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | Sep 3, 2026 | 1.05M; standard rates up to 272K input | $10 / $50 | No. ChatGPT Plus ($20) and up, not on Free or Go | The hardest reasoning, coding and research; five effort levels up to max. Checked Sep 2026. |
| GPT-5.6 Sol | OpenAI | Jul 9, 2026 | 1.05M; standard rates up to 272K | $4 / $20 promotional, held at least through Nov 21, 2026 | No. ChatGPT Plus and up | Daily professional work with the reasoning slider. Note the long-context band doubles input above 272K. Checked Sep 2026. |
| GPT-5.6 Terra | OpenAI | Jul 9, 2026 | 1.05M | $2 / $12 | Limited, in ChatGPT Work and Codex on the desktop app; full on Plus | Mid-priced agents and coding; Perplexity uses it for Computer subagents. Checked Sep 2026. |
| GPT-5.6 Luna | OpenAI | Jul 9, 2026; default on Free since Aug 6 | 1.05M on the API; 27K in the free ChatGPT app | $0.20 / $1.20 | Yes. Unlimited text chat on ChatGPT Free | Everyday chat and high-volume automation. Checked Sep 2026. |
| Claude Fable 5.1 | Anthropic | Sep 1, 2026 | 1M; 128K output | $10 / $50; cache reads $0.25 | No. Usage credits on Claude Pro, half of weekly limits on Max | Long-running agents and demanding reasoning; adaptive thinking is always on. Checked Sep 2026. |
| Claude Opus 5 | Anthropic | Jul 24, 2026 | 1M; 128K output | $5 / $25 | No. Claude Pro ($20) and up | Anthropic's recommended default: agentic coding and enterprise work at half Fable's price. Checked Sep 2026. |
| Claude Sonnet 5 | Anthropic | Jun 2026; default on Free and Pro since Jun 30 | 1M; 128K output | $2 / $10 | Yes. Claude Free | Speed with near-frontier quality; the best writing model you can use free. Checked Sep 2026. |
| Claude Haiku 4.5 | Anthropic | Oct 2025 | 200K; 64K output | $1 / $5 | Yes. Claude Free | High-volume, low-latency tasks; retirement not before Oct 15, 2026. Checked Sep 2026. |
| Gemini 3.1 Pro (preview) | Feb 19, 2026 | 1,048,576 | $2 / $12 up to 200K; $4 / $18 above | Varying access in the free Gemini app; expanded on AI Pro | Multimodal reasoning over text, audio, video and whole repositories. Still Google's Pro model in Sep 2026; 3.5 Pro was in partner testing as of Jul 21. | |
| Gemini 3.8 Flash | Sep 2, 2026 | About 1M; not stated on the pricing page, same family as 3.6 Flash | $0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50 | Free API tier yes; in the app only on AI Pro and Ultra | Long-horizon software engineering and agents at Flash prices. Checked Sep 2026. | |
| Gemini 3.6 Flash | Jul 21, 2026 | 1M; 64K output | $0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50 | Yes. Default model in the free Gemini app | The everyday workhorse; 17% fewer output tokens than 3.5 Flash. Checked Sep 2026. | |
| Muse Spark 1.3 | Meta | Sep 2, 2026 | 1M | $1.25 / $4.25; $0.10 / $0.20 on the contributor endpoint that lets Meta train on your traffic | Meta AI app and meta.ai; the max variant was in limited partner preview at launch | Cheap frontier-class multimodal agents. Pricing from Meta's developer pages via independent trackers, Sep 2026. |
| Muse Spark 1.1 | Meta | Jul 8, 2026 | 1M | $1.25 / $4.25 | Yes. Thinking mode in the Meta AI app and on meta.ai | Multi-agent orchestration and computer use; OpenAI- and Anthropic-compatible API. Checked Sep 2026. |
| Grok 4.7 | xAI | Sep 21, 2026 | 500K; rates double at 200K | $2 / $6; cached $0.50 | Free access in Grok Build; grok.com's free tier runs lighter models | Coding and knowledge work at a low price; four effort levels. Checked launch day, Sep 2026. |
| Grok 4.3 | xAI | Apr 30, 2026 | 1M | $1.25 / $2.50 | Not on the free tier | Long-context general work on a budget. Checked Sep 2026. |
| Mistral Medium 3.5 | Mistral | May 22, 2026 | 256K | $1.50 / $7.50 | Yes. Default model in the free Le Chat and Vibe tiers | Coding agents and Work mode; 128B dense open weights under a modified MIT license. Checked Sep 2026. |
| Mistral Large 3 | Mistral | Dec 2, 2025 | 256K | $0.50 / $1.50 | Yes. Le Chat free | Cheapest open-weight frontier-class model; Apache 2.0, 675B total, 41B active. Checked Sep 2026. |
| Mistral Small 4 | Mistral | Mar 16, 2026 | 256K | €0.12 / €0.50 (Mistral lists it in euros) | Yes. Le Chat free, and it runs locally | Reasoning, vision and coding in one small Apache 2.0 model. Checked Sep 2026. |
| DeepSeek V4 Pro | DeepSeek | Preview Apr 24, 2026; GA Aug 13, 2026 | 1M; 384K output | $1.32 / $3.96 peak; $0.66 / $1.98 off-peak | Yes. Expert Mode in the free DeepSeek app and web | Open-weight reasoning at a fraction of Western prices; 1.6T total, 49B active. Checked Sep 2026. |
| DeepSeek V4.1 Flash | DeepSeek | V4 Flash Apr 24, 2026; now served as V4.1 | 1M; 384K output | $0.30 / $1.20 peak; $0.15 / $0.60 off-peak | Yes. Default in the free DeepSeek app | The cheapest capable model on this list; weekends bill at off-peak. Checked Sep 2026. |
| Qwen3.8-Max (0902) | Alibaba | GA Aug 2026; 0902 snapshot Sep 2, 2026 | 1M | $2 / $6 international list | Yes. Qwen Chat is free | Engineering-scale coding, tool use, images and up to two hours of video input. Checked Sep 2026. |
| Qwen3.7-Plus | Alibaba | May 26, 2026 | 256K | CNY 3 / CNY 12 international list, about $0.40 / $1.65 | Yes. Qwen Chat | Budget workhorse; Model Studio was showing a limited-time 20% discount in Sep 2026. |
How to read the table
Input is what you send, output is what the model writes back, and every lab on this list charges more for output, usually four to six times more. A long answer costs more than a long question. Three modifiers change the sticker price on almost every row. Prompt caching cuts repeat input to a tenth or less of the base rate, which matters if you send the same system prompt or document over and over. Batch processing, where the answer can wait up to a day, halves the bill at OpenAI, Anthropic, Google and Alibaba. And long-context bands raise the rate for the whole request once the prompt crosses a line: 272K tokens at OpenAI, 200K at Google and xAI. DeepSeek uses time instead, with off-peak rates 50% below peak outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.
Context is the total the model can hold in one request, in tokens. A million tokens is roughly 555,000 words on Anthropic's current tokenizer and about 750,000 on older ones, so the 1M rows can read a long novel or a mid-sized codebase in one go. The consumer apps often expose less than the API: the free ChatGPT tier holds 27K tokens for instant answers, the free Gemini app 32K. Where a row says a model is free for consumers, it means you can use it in the maker's own chat app without paying, not that the API is free. Google's API is the exception, with a real free tier that uses your prompts to improve its products.
Two rows carry a caveat. Gemini 3.8 Flash's context window is not stated on Google's pricing page; we list about 1M because it is built on the 3.6 Flash family whose model card says 1M. Meta's Muse Spark 1.3 prices come from Meta's developer pricing page as reported by three independent trackers in the first week of September 2026, since Meta's own page did not render for us; the 1.1 price of $1.25/$4.25 is confirmed on Meta's developer blog and 1.3 kept it.
What changed since summer 2026
Since June the top of the market has turned over completely. OpenAI made the GPT-5.6 family generally available on July 9, cut its API prices on July 30 and again on August 21, made Luna the unlimited free model on August 6, then shipped GPT-6 Astra on September 3 as its first model designated Critical for cybersecurity under its Preparedness Framework. Anthropic released Sonnet 5 in June, Opus 5 on July 24 at half the price of Fable 5, and Fable 5.1 on September 1, which kept the $10/$50 rate but cut cache reads to $0.25. Astra and Fable 5.1 now match to the cent on every published line except that one.
Google went the other way, competing on price. Gemini 3.6 Flash arrived July 21 and 3.8 Flash on September 2, both at an introductory $0.75/$3.75 that doubles on January 1, 2027. Gemini 3.5 Pro, which Google said was in partner testing in July, had still not replaced 3.1 Pro on the pricing page in September. Meta finished its pivot away from open-weight Llama: Muse Spark, its first closed model from Meta Superintelligence Labs, got a paid API on July 8 and reached version 1.3 on September 2, while an open-weight small model, Muse Glimmer, shipped on August 10 under Apache 2.0.
The Chinese labs kept the floor low. DeepSeek V4 Pro went GA on August 13 with the new peak and off-peak pricing three days later; Qwen3.8-Max went GA in August and got a stronger 0902 snapshot on September 2. xAI shipped Grok 4.6 in August and, after several delays, Grok 4.7 on September 21 at the same $2/$6. Mistral's newest flagship remains Medium 3.5 from May 22, an open-weight 128B model that replaced its coding and reasoning lines in Le Chat.
Which model to pick for which job
| Job | First choice | Cheaper alternative |
|---|---|---|
| Hardest reasoning, research, agents that run for hours | Claude Fable 5.1 or GPT-6 Astra | Claude Opus 5 at half the price |
| Everyday coding in an editor | Claude Opus 5 or GPT-5.6 Sol | Gemini 3.8 Flash or Grok 4.7 |
| Writing and editing | Claude Sonnet 5 | GPT-5.6 Terra |
| Long documents, video, audio | Gemini 3.1 Pro | Gemini 3.6 Flash |
| High-volume classification and extraction | GPT-5.6 Luna or Claude Haiku 4.5 | DeepSeek V4.1 Flash off-peak |
| Running locally or on your own servers | Mistral Medium 3.5 or DeepSeek V4 Pro (open weights) | Mistral Small 4 |
| Lowest possible bill, data privacy not critical | Muse Spark 1.3 contributor tier | DeepSeek V4.1 Flash |
For the chat apps rather than the API, the choice collapses to which assistant you already like: our chatbot comparison maps these models to the plans that include them. Developers who route between models from an editor should read the Cursor card; almost every model above is selectable there within days of release.
What to do: default to the mid-tier, escalate only on evidence.
Start with Sonnet 5, GPT-5.6 Terra or Gemini 3.8 Flash for any new job. Move up to Opus 5, Sol or 3.1 Pro only when your own tests fail at the mid-tier, and reserve Fable 5.1 and Astra for the work that pays for itself. If you are price-sensitive and the data is not sensitive, DeepSeek and Qwen are now credible at a tenth of the cost. Recheck this page monthly; six of the twenty-two rows changed in the four weeks before publication.
Quick answers
What is the best AI model right now?
As of September 2026 the two frontier flagships are GPT-6 Astra and Claude Fable 5.1, both priced at $10 per million input tokens and $50 per million output. Claude Opus 5 and Gemini 3.1 Pro sit just below at $5/$25 and $2/$12. For most everyday jobs a mid-tier model such as Sonnet 5, GPT-5.6 Terra or Gemini 3.8 Flash is the better buy.
What is the cheapest capable AI model?
DeepSeek's V4.1 Flash at $0.15 per million input tokens and $0.60 output during off-peak hours, followed by GPT-5.6 Luna at $0.20/$1.20 and Meta's Muse Spark 1.3 contributor tier at $0.10/$0.20 if you allow Meta to train on your traffic. Gemini 3.8 Flash at $0.75/$3.75 is the cheapest of the Western frontier-lab workhorses until its introductory price ends on December 31, 2026.
Which AI models can I use free?
GPT-5.6 Luna is unlimited on ChatGPT Free, Claude Sonnet 5 and Haiku 4.5 are on Claude Free, Gemini 3.6 Flash and some Gemini 3.1 Pro are in the free Gemini app, Muse Spark is in the free Meta AI app, DeepSeek V4 Pro is in the free DeepSeek app, Qwen3.8-Max is in the free Qwen Chat, and Mistral's models are in the free Le Chat tier.
Sources
- OpenAI — API pricing for GPT-6 Astra and the GPT-5.6 family, checked September 2026
- Anthropic — Claude models overview with context windows and pricing, checked September 2026
- Google — Gemini Developer API pricing, checked September 2026
- xAI — Grok models and pricing, checked September 2026
- DeepSeek — Models and pricing with peak and off-peak rates, checked September 2026
- Alibaba Cloud Model Studio — Model inference pricing, checked September 2026


