AI model tracker, September 2026

Eight labs, twenty-two current models, one table. Every row was checked against the maker's own documentation or pricing page in September 2026; where a figure could not be verified we say so or left the row out. Prices are per million tokens on the standard API tier, US dollars unless marked.

A neat grid of small glossy playing cards laid face up on a dark navy table, each card a different solid color, lit from one side
AI-generated illustration

Figures come from vendor pricing pages, model cards and release notes read in September 2026, cross-checked against independent price trackers. Some vendor links on this site are affiliate links; they never affect what we list. See our disclosure.

The table

ModelMakerReleasedContextAPI in / out per 1MFree for consumersBest for
GPT-6 AstraOpenAISep 3, 20261.05M; standard rates up to 272K input$10 / $50No. ChatGPT Plus ($20) and up, not on Free or GoThe hardest reasoning, coding and research; five effort levels up to max. Checked Sep 2026.
GPT-5.6 SolOpenAIJul 9, 20261.05M; standard rates up to 272K$4 / $20 promotional, held at least through Nov 21, 2026No. ChatGPT Plus and upDaily professional work with the reasoning slider. Note the long-context band doubles input above 272K. Checked Sep 2026.
GPT-5.6 TerraOpenAIJul 9, 20261.05M$2 / $12Limited, in ChatGPT Work and Codex on the desktop app; full on PlusMid-priced agents and coding; Perplexity uses it for Computer subagents. Checked Sep 2026.
GPT-5.6 LunaOpenAIJul 9, 2026; default on Free since Aug 61.05M on the API; 27K in the free ChatGPT app$0.20 / $1.20Yes. Unlimited text chat on ChatGPT FreeEveryday chat and high-volume automation. Checked Sep 2026.
Claude Fable 5.1AnthropicSep 1, 20261M; 128K output$10 / $50; cache reads $0.25No. Usage credits on Claude Pro, half of weekly limits on MaxLong-running agents and demanding reasoning; adaptive thinking is always on. Checked Sep 2026.
Claude Opus 5AnthropicJul 24, 20261M; 128K output$5 / $25No. Claude Pro ($20) and upAnthropic's recommended default: agentic coding and enterprise work at half Fable's price. Checked Sep 2026.
Claude Sonnet 5AnthropicJun 2026; default on Free and Pro since Jun 301M; 128K output$2 / $10Yes. Claude FreeSpeed with near-frontier quality; the best writing model you can use free. Checked Sep 2026.
Claude Haiku 4.5AnthropicOct 2025200K; 64K output$1 / $5Yes. Claude FreeHigh-volume, low-latency tasks; retirement not before Oct 15, 2026. Checked Sep 2026.
Gemini 3.1 Pro (preview)GoogleFeb 19, 20261,048,576$2 / $12 up to 200K; $4 / $18 aboveVarying access in the free Gemini app; expanded on AI ProMultimodal reasoning over text, audio, video and whole repositories. Still Google's Pro model in Sep 2026; 3.5 Pro was in partner testing as of Jul 21.
Gemini 3.8 FlashGoogleSep 2, 2026About 1M; not stated on the pricing page, same family as 3.6 Flash$0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50Free API tier yes; in the app only on AI Pro and UltraLong-horizon software engineering and agents at Flash prices. Checked Sep 2026.
Gemini 3.6 FlashGoogleJul 21, 20261M; 64K output$0.75 / $3.75 through Dec 31, 2026, then $1.50 / $7.50Yes. Default model in the free Gemini appThe everyday workhorse; 17% fewer output tokens than 3.5 Flash. Checked Sep 2026.
Muse Spark 1.3MetaSep 2, 20261M$1.25 / $4.25; $0.10 / $0.20 on the contributor endpoint that lets Meta train on your trafficMeta AI app and meta.ai; the max variant was in limited partner preview at launchCheap frontier-class multimodal agents. Pricing from Meta's developer pages via independent trackers, Sep 2026.
Muse Spark 1.1MetaJul 8, 20261M$1.25 / $4.25Yes. Thinking mode in the Meta AI app and on meta.aiMulti-agent orchestration and computer use; OpenAI- and Anthropic-compatible API. Checked Sep 2026.
Grok 4.7xAISep 21, 2026500K; rates double at 200K$2 / $6; cached $0.50Free access in Grok Build; grok.com's free tier runs lighter modelsCoding and knowledge work at a low price; four effort levels. Checked launch day, Sep 2026.
Grok 4.3xAIApr 30, 20261M$1.25 / $2.50Not on the free tierLong-context general work on a budget. Checked Sep 2026.
Mistral Medium 3.5MistralMay 22, 2026256K$1.50 / $7.50Yes. Default model in the free Le Chat and Vibe tiersCoding agents and Work mode; 128B dense open weights under a modified MIT license. Checked Sep 2026.
Mistral Large 3MistralDec 2, 2025256K$0.50 / $1.50Yes. Le Chat freeCheapest open-weight frontier-class model; Apache 2.0, 675B total, 41B active. Checked Sep 2026.
Mistral Small 4MistralMar 16, 2026256K€0.12 / €0.50 (Mistral lists it in euros)Yes. Le Chat free, and it runs locallyReasoning, vision and coding in one small Apache 2.0 model. Checked Sep 2026.
DeepSeek V4 ProDeepSeekPreview Apr 24, 2026; GA Aug 13, 20261M; 384K output$1.32 / $3.96 peak; $0.66 / $1.98 off-peakYes. Expert Mode in the free DeepSeek app and webOpen-weight reasoning at a fraction of Western prices; 1.6T total, 49B active. Checked Sep 2026.
DeepSeek V4.1 FlashDeepSeekV4 Flash Apr 24, 2026; now served as V4.11M; 384K output$0.30 / $1.20 peak; $0.15 / $0.60 off-peakYes. Default in the free DeepSeek appThe cheapest capable model on this list; weekends bill at off-peak. Checked Sep 2026.
Qwen3.8-Max (0902)AlibabaGA Aug 2026; 0902 snapshot Sep 2, 20261M$2 / $6 international listYes. Qwen Chat is freeEngineering-scale coding, tool use, images and up to two hours of video input. Checked Sep 2026.
Qwen3.7-PlusAlibabaMay 26, 2026256KCNY 3 / CNY 12 international list, about $0.40 / $1.65Yes. Qwen ChatBudget workhorse; Model Studio was showing a limited-time 20% discount in Sep 2026.

How to read the table

Input is what you send, output is what the model writes back, and every lab on this list charges more for output, usually four to six times more. A long answer costs more than a long question. Three modifiers change the sticker price on almost every row. Prompt caching cuts repeat input to a tenth or less of the base rate, which matters if you send the same system prompt or document over and over. Batch processing, where the answer can wait up to a day, halves the bill at OpenAI, Anthropic, Google and Alibaba. And long-context bands raise the rate for the whole request once the prompt crosses a line: 272K tokens at OpenAI, 200K at Google and xAI. DeepSeek uses time instead, with off-peak rates 50% below peak outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

Context is the total the model can hold in one request, in tokens. A million tokens is roughly 555,000 words on Anthropic's current tokenizer and about 750,000 on older ones, so the 1M rows can read a long novel or a mid-sized codebase in one go. The consumer apps often expose less than the API: the free ChatGPT tier holds 27K tokens for instant answers, the free Gemini app 32K. Where a row says a model is free for consumers, it means you can use it in the maker's own chat app without paying, not that the API is free. Google's API is the exception, with a real free tier that uses your prompts to improve its products.

Two rows carry a caveat. Gemini 3.8 Flash's context window is not stated on Google's pricing page; we list about 1M because it is built on the 3.6 Flash family whose model card says 1M. Meta's Muse Spark 1.3 prices come from Meta's developer pricing page as reported by three independent trackers in the first week of September 2026, since Meta's own page did not render for us; the 1.1 price of $1.25/$4.25 is confirmed on Meta's developer blog and 1.3 kept it.

What changed since summer 2026

Since June the top of the market has turned over completely. OpenAI made the GPT-5.6 family generally available on July 9, cut its API prices on July 30 and again on August 21, made Luna the unlimited free model on August 6, then shipped GPT-6 Astra on September 3 as its first model designated Critical for cybersecurity under its Preparedness Framework. Anthropic released Sonnet 5 in June, Opus 5 on July 24 at half the price of Fable 5, and Fable 5.1 on September 1, which kept the $10/$50 rate but cut cache reads to $0.25. Astra and Fable 5.1 now match to the cent on every published line except that one.

Google went the other way, competing on price. Gemini 3.6 Flash arrived July 21 and 3.8 Flash on September 2, both at an introductory $0.75/$3.75 that doubles on January 1, 2027. Gemini 3.5 Pro, which Google said was in partner testing in July, had still not replaced 3.1 Pro on the pricing page in September. Meta finished its pivot away from open-weight Llama: Muse Spark, its first closed model from Meta Superintelligence Labs, got a paid API on July 8 and reached version 1.3 on September 2, while an open-weight small model, Muse Glimmer, shipped on August 10 under Apache 2.0.

The Chinese labs kept the floor low. DeepSeek V4 Pro went GA on August 13 with the new peak and off-peak pricing three days later; Qwen3.8-Max went GA in August and got a stronger 0902 snapshot on September 2. xAI shipped Grok 4.6 in August and, after several delays, Grok 4.7 on September 21 at the same $2/$6. Mistral's newest flagship remains Medium 3.5 from May 22, an open-weight 128B model that replaced its coding and reasoning lines in Le Chat.

Which model to pick for which job

JobFirst choiceCheaper alternative
Hardest reasoning, research, agents that run for hoursClaude Fable 5.1 or GPT-6 AstraClaude Opus 5 at half the price
Everyday coding in an editorClaude Opus 5 or GPT-5.6 SolGemini 3.8 Flash or Grok 4.7
Writing and editingClaude Sonnet 5GPT-5.6 Terra
Long documents, video, audioGemini 3.1 ProGemini 3.6 Flash
High-volume classification and extractionGPT-5.6 Luna or Claude Haiku 4.5DeepSeek V4.1 Flash off-peak
Running locally or on your own serversMistral Medium 3.5 or DeepSeek V4 Pro (open weights)Mistral Small 4
Lowest possible bill, data privacy not criticalMuse Spark 1.3 contributor tierDeepSeek V4.1 Flash

For the chat apps rather than the API, the choice collapses to which assistant you already like: our chatbot comparison maps these models to the plans that include them. Developers who route between models from an editor should read the Cursor card; almost every model above is selectable there within days of release.

What to do: default to the mid-tier, escalate only on evidence.

Start with Sonnet 5, GPT-5.6 Terra or Gemini 3.8 Flash for any new job. Move up to Opus 5, Sol or 3.1 Pro only when your own tests fail at the mid-tier, and reserve Fable 5.1 and Astra for the work that pays for itself. If you are price-sensitive and the data is not sensitive, DeepSeek and Qwen are now credible at a tenth of the cost. Recheck this page monthly; six of the twenty-two rows changed in the four weeks before publication.

Quick answers

What is the best AI model right now?

As of September 2026 the two frontier flagships are GPT-6 Astra and Claude Fable 5.1, both priced at $10 per million input tokens and $50 per million output. Claude Opus 5 and Gemini 3.1 Pro sit just below at $5/$25 and $2/$12. For most everyday jobs a mid-tier model such as Sonnet 5, GPT-5.6 Terra or Gemini 3.8 Flash is the better buy.

What is the cheapest capable AI model?

DeepSeek's V4.1 Flash at $0.15 per million input tokens and $0.60 output during off-peak hours, followed by GPT-5.6 Luna at $0.20/$1.20 and Meta's Muse Spark 1.3 contributor tier at $0.10/$0.20 if you allow Meta to train on your traffic. Gemini 3.8 Flash at $0.75/$3.75 is the cheapest of the Western frontier-lab workhorses until its introductory price ends on December 31, 2026.

Which AI models can I use free?

GPT-5.6 Luna is unlimited on ChatGPT Free, Claude Sonnet 5 and Haiku 4.5 are on Claude Free, Gemini 3.6 Flash and some Gemini 3.1 Pro are in the free Gemini app, Muse Spark is in the free Meta AI app, DeepSeek V4 Pro is in the free DeepSeek app, Qwen3.8-Max is in the free Qwen Chat, and Mistral's models are in the free Le Chat tier.