Research

Crown Chasers and Price Breakers

The game theory of the AI labs.

Astra. Fable. Grok. The names sound expensive. The engineering question is whether the extra intelligence earns its bill.

I think of the labs as players in a repeated game, with two strategies worth watching. Crown chasers push for the highest capability they can demonstrate. Their flagship has to do something the competition cannot. Price breakers attack the cost of capable models. Their threat is a model good enough to make an expensive incumbent uncomfortable.

These are competitive strategies. One lab can play both. A flagship can chase the crown while its little brother attacks the bill.

When the little brother makes a task cheaper, the flagship has to defend its premium. A lab might cut the price, push capability further (like loosening guardrails to widen access to capabilities), or win on something the cheaper model cannot do. Or win on more reliable tool use or lower latency. Every move changes what the next player has to beat.

My bet is that this pressure is a powerful driver of progress. After reviewing the data on the Frontier class and the Pareto class, I can observe capability at the Pareto class that did not exist in March of this year. Pareto models are getting better per dollar.

Frontier models and Pareto models

Crown chasers and price breakers describe how the labs compete. Frontier and Pareto describe what we buy. Here is the vocabulary I use.

The frontier is a moving capability boundary. The Frontier Model Forum uses the term broadly for the current state of the art. I look for evidence across difficult reasoning, engineering, and tool use. A model can keep every capability it has today and lose its frontier position when a rival advances. The unlock with a frontier model is the degree to which it can challenge or improve a technical decision, or contribute an insight for the greater whole. Call it higher-order thinking: reasoning through the consequences of a decision and navigating the branches of possibility.

Our Pareto definition adds a selection rule to Pareto efficiency. Pareto efficiency identifies tradeoffs: improving one objective means giving up another. Our workload supplies the quality bar and the evaluation meter. Evaluations establish which configurations clear the bar. Among those, we choose the lowest total cost, and retries and human repair count in that bill.

I run such selections by starting at the frontier and testing downwards. The strongest model establishes what a good result looks like. Each cheaper model then tests how well I have specified the work and equipped the agent.

A failure at a cheaper model can mean two things. The model is too weak, or the workload is underspecified: vague instructions, missing context, a brittle environment. A frontier model tends to reason its way around those defects, at a cost, and hides them. A cheaper model exposes them. Fixing what it exposes improves the whole system, so testing downwards is a quality step as well as a cost step.

These labels can overlap. A frontier model may also be Pareto-efficient, and may be our cheapest reliable choice, all depending on the nature of the task. Model family, size, and open or closed weights do not settle that question. Neither does a low token price. The comparison belongs to a configuration doing a job. Anthropic's model-selection guide makes the same practical case for testing capability, speed, cost, and reasoning effort against actual work. Anthropic presents two recommended paths: efficiency-first, or capability-first. I believe that starting capability-first, then progressing downwards, is the more generally applicable choice. It lets you find the specific definition of the system more closely than starting bottom up.

The good news is that the number of models delivering near-frontier intelligence at Pareto costs is growing. Comparing early September with March shows how quickly the truly competent Pareto-class candidates have multiplied.

March: the challenger box is empty

Take the Artificial Analysis leaderboard captured on 6 March 2026. Among 138 configurations with a measured score and usable cost, Gemini 3.1 Pro Preview held the highest score.

The challenger box sets two conditions for taking on the leader:

  • Score: at least 80% of the leader’s score.
  • Cost: at most 20% of the leader’s bill for running the same index.

Meet both, and a configuration is an 80/20 challenger. Near the leading score. A fraction of the bill.

In March, nobody qualified.

The challenger box is empty, 6 March 20260 80/20 challengers among 138 eligible configurations. Shading marks the challenger box. The staircase is the Pareto frontier. The cost axis is logarithmic. Hover or tap a point for its model. With the chart focused, use the left and right arrow keys to inspect points.Score relative to the leader (%)0204060801000.1%1%10%100%1000%Phi-4. index: 10.4. index cost: $4.27.Granite 4.0 H Small. index: 10.8. index cost: $4.48.Gemma 3n E4B Instruct. index: 6.4. index cost: $5.72.Nova Micro. index: 10.3. index cost: $6.38.gpt-oss-20B (low). index: 20.8. index cost: $7.68.ERNIE 4.5 300B A47B. index: 15.0. index cost: $8.09.LFM2 24B A2B. index: 10.5. index cost: $8.78.NVIDIA Nemotron Nano 9B V2 (Reasoning). index: 14.8. index cost: $10.53.Llama 3.2 Instruct 11B (Vision). index: 8.7. index cost: $12.67.NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning). index: 13.2. index cost: $13.73.Qwen3 1.7B (Non-reasoning). index: 6.8. index cost: $15.77.gpt-oss-120B (low). index: 24.5. index cost: $15.90.Ministral 3 3B. index: 11.2. index cost: $15.93.Ling-flash-2.0. index: 15.7. index cost: $16.73.Olmo 3 7B Think. index: 9.4. index cost: $17.51.Ling-mini-2.0. index: 9.2. index cost: $18.41.Llama Nemotron Super 49B v1.5 (Non-reasoning). index: 14.6. index cost: $20.69.gpt-oss-20B (high). index: 24.5. index cost: $20.69.Qwen3 0.6B (Non-reasoning). index: 5.7. index cost: $21.34.Grok 4.1 Fast (Non-reasoning). index: 23.6. index cost: $21.37.MiMo-V2-Flash (Non-reasoning). index: 30.4. index cost: $21.38.Ministral 3 8B. index: 14.8. index cost: $21.45.Mistral Small 3.2. index: 15.1. index cost: $21.97.GLM-4.7-Flash (Non-reasoning). index: 22.1. index cost: $22.19.Ministral 3 14B. index: 16.0. index cost: $23.19.Llama 4 Scout. index: 13.5. index cost: $27.06.Olmo 3 7B Instruct. index: 8.2. index cost: $29.68.Hermes 4 - Llama-3.1 70B (Non-reasoning). index: 12.6. index cost: $29.85.Qwen3 0.6B (Reasoning). index: 6.5. index cost: $30.42.Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning). index: 19.4. index cost: $33.39.Llama Nemotron Super 49B v1.5 (Reasoning). index: 18.7. index cost: $37.18.Reka Flash 3. index: 9.5. index cost: $37.30.Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning). index: 21.6. index cost: $38.21.Mistral Large 3. index: 22.8. index cost: $38.67.Grok 4.1 Fast (Reasoning). index: 38.6. index cost: $39.62.GLM-4.7-Flash (Reasoning). index: 30.1. index cost: $40.01.NVIDIA Nemotron 3 Nano 30B A3B (Reasoning). index: 24.3. index cost: $45.90.Hermes 4 - Llama-3.1 70B (Reasoning). index: 16.0. index cost: $48.10.Mistral Medium 3.1. index: 21.3. index cost: $51.65.GPT-5 nano (high). index: 26.8. index cost: $51.71.Qwen3 1.7B (Reasoning). index: 8.0. index cost: $52.53.Llama 4 Maverick. index: 18.4. index cost: $55.96.Ring-flash-2.0. index: 14.0. index cost: $59.81.GLM-4.6V (Non-reasoning). index: 17.1. index cost: $61.16.Gemini 3 Flash Preview (Non-reasoning). index: 35.0. index cost: $65.98.gpt-oss-120B (high). index: 33.3. index cost: $67.37.MiMo-V2-Flash (Feb 2026). index: 41.5. index cost: $68.26.Olmo 3.1 32B Instruct. index: 12.2. index cost: $70.48.DeepSeek V3.2 (Reasoning). index: 41.7. index cost: $70.64.Qwen3 VL 8B Instruct. index: 14.3. index cost: $72.73.Qwen3 Coder 30B A3B Instruct. index: 20.0. index cost: $72.84.Nova 2.0 Lite (low). index: 24.6. index cost: $72.96.Llama 3.1 Nemotron Instruct 70B. index: 13.4. index cost: $73.54.Step 3.5 Flash. index: 37.8. index cost: $74.80.Qwen3 30B A3B 2507 Instruct. index: 15.0. index cost: $75.44.Nova 2.0 Omni (low). index: 23.2. index cost: $75.63.KAT-Coder-Pro V1. index: 36.0. index cost: $76.16.NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning). index: 10.1. index cost: $78.55.Mercury 2. index: 32.8. index cost: $80.43.Qwen3 Coder Next. index: 28.3. index cost: $84.56.Qwen3 Omni 30B A3B Instruct. index: 10.7. index cost: $85.07.Llama 3.3 Instruct 70B. index: 14.5. index cost: $86.26.Magistral Small 1.2. index: 18.2. index cost: $89.83.Gemini 3.1 Flash-Lite Preview. index: 33.5. index cost: $93.60.DeepSeek V3.2 (Non-reasoning). index: 32.1. index cost: $103.16.Grok 3 mini Reasoning (high). index: 32.1. index cost: $104.77.Nova 2.0 Omni (medium). index: 28.0. index cost: $109.45.Grok Code Fast 1. index: 28.7. index cost: $116.70.Qwen3 VL 30B A3B Instruct. index: 16.1. index cost: $120.75.NVIDIA Nemotron Nano 12B v2 VL (Reasoning). index: 14.9. index cost: $123.08.Qwen3.5 27B (Non-reasoning). index: 37.2. index cost: $123.35.MiniMax-M2.5. index: 41.9. index cost: $124.58.Qwen3.5 35B A3B (Non-reasoning). index: 30.7. index cost: $125.61.Qwen3 Omni 30B A3B (Reasoning). index: 15.6. index cost: $132.31.Qwen3 Next 80B A3B Instruct. index: 20.1. index cost: $132.97.EXAONE 4.0 32B (Non-reasoning). index: 11.7. index cost: $138.74.Qwen3 30B A3B 2507 (Reasoning). index: 22.4. index cost: $140.96.Kimi K2.5 (Non-reasoning). index: 37.3. index cost: $141.12.Qwen3 VL 8B (Reasoning). index: 16.7. index cost: $148.48.Qwen3 VL 30B A3B (Reasoning). index: 19.7. index cost: $156.75.EXAONE 4.0 32B (Reasoning). index: 16.7. index cost: $163.15.Qwen3.5 122B A10B (Non-reasoning). index: 35.9. index cost: $166.37.GPT-5 mini (high). index: 41.2. index cost: $168.46.Hermes 4 - Llama-3.1 405B (Non-reasoning). index: 17.6. index cost: $178.70.Nova 2.0 Lite (medium). index: 29.7. index cost: $182.88.GLM-4.6V (Reasoning). index: 23.4. index cost: $185.47.Qwen3.5 397B A17B (Non-reasoning). index: 40.1. index cost: $186.37.Qwen3 Max. index: 31.4. index cost: $196.60.GPT-5.1 Codex mini (high). index: 38.6. index cost: $201.60.Nova 2.0 Pro Preview (low). index: 31.9. index cost: $204.69.Claude 4.5 Haiku (Non-reasoning). index: 31.1. index cost: $205.11.GPT-5.2 (Non-reasoning). index: 33.6. index cost: $225.44.GLM-5 (Non-reasoning). index: 40.6. index cost: $239.99.Qwen3 235B A22B 2507 Instruct. index: 25.0. index cost: $255.94.Qwen3 VL 235B A22B Instruct. index: 20.8. index cost: $264.19.Qwen3 VL 32B Instruct. index: 17.2. index cost: $275.36.Gemini 3 Flash Preview (Reasoning). index: 46.4. index cost: $278.26.Qwen3 Max Thinking (Preview). index: 32.5. index cost: $293.35.Hermes 4 - Llama-3.1 405B (Reasoning). index: 18.6. index cost: $294.45.Command A. index: 13.5. index cost: $299.18.Qwen3.5 27B (Reasoning). index: 42.1. index cost: $299.26.Qwen3.5 35B A3B (Reasoning). index: 37.1. index cost: $301.70.DeepSeek R1 0528 (May '25). index: 27.1. index cost: $339.08.Qwen3.5 122B A10B (Reasoning). index: 41.6. index cost: $353.94.Gemini 3 Pro Preview (low). index: 41.3. index cost: $354.97.Qwen3 Next 80B A3B (Reasoning). index: 26.7. index cost: $357.34.Kimi K2.5 (Reasoning). index: 46.8. index cost: $370.66.Nova Premier. index: 19.0. index cost: $375.98.Magistral Medium 1.2. index: 27.1. index cost: $387.10.Nova 2.0 Lite (Non-reasoning). index: 18.0. index cost: $411.01.Qwen3.5 397B A17B (Reasoning). index: 45.0. index cost: $417.54.Qwen3 VL 235B A22B (Reasoning). index: 27.6. index cost: $427.93.Nova 2.0 Pro Preview (medium). index: 35.7. index cost: $466.55.Llama 3.1 Nemotron Ultra 253B v1 (Reasoning). index: 15.0. index cost: $486.79.GLM-5 (Reasoning). index: 49.8. index cost: $547.04.Claude Sonnet 4.6 (Non-reasoning, Low Effort). index: 42.6. index cost: $554.48.Qwen3 235B A22B 2507 (Reasoning). index: 29.5. index cost: $569.51.Claude 4.5 Haiku (Reasoning). index: 37.1. index cost: $583.03.Nova 2.0 Omni (Non-reasoning). index: 16.6. index cost: $598.35.Gemini 2.5 Pro. index: 34.6. index cost: $648.32.Qwen3 Max Thinking. index: 39.9. index cost: $668.54.GPT-5.2 (medium). index: 46.6. index cost: $699.65.Qwen3 Coder 480B A35B Instruct. index: 24.8. index cost: $779.13.Gemini 3.1 Pro Preview. index: 57.2. index cost: $892.28.Jamba 1.7 Large. index: 10.9. index cost: $965.33.Qwen3 VL 32B (Reasoning). index: 24.7. index cost: $971.59.o3. index: 38.4. index cost: $1,024.79.Llama 3.1 Instruct 405B. index: 17.4. index cost: $1,219.16.Nova 2.0 Pro Preview (Non-reasoning). index: 23.1. index cost: $1,300.10.Claude Sonnet 4.6 (Non-reasoning, High Effort). index: 44.4. index cost: $1,396.63.Claude Opus 4.6 (Non-reasoning, High Effort). index: 46.5. index cost: $1,451.04.Grok 4. index: 41.5. index cost: $1,568.34.GPT-5.3 Codex (xhigh). index: 54.0. index cost: $1,653.80.GPT-5.2 (xhigh). index: 51.3. index cost: $2,304.22.GPT-5.4 (xhigh). index: 57.0. index cost: $2,950.85.GPT-5.2 Codex (xhigh). index: 49.0. index cost: $3,244.33.Claude Sonnet 4.6 (max). index: 51.7. index cost: $3,959.36.Claude Opus 4.6 (max). index: 53.0. index cost: $4,969.68.123Cost relative to the leader (%)

0 80/20 challengers / 138 eligible configurations. Shading marks the challenger box. The staircase is the Pareto frontier. The cost axis is logarithmic.

  1. Gemini 3.1 Pro Preview. Leading score. 57.2 index, $892.28 index cost.
  2. Gemini 3 Flash Preview (Reasoning). Cheapest efficient configuration at 80% of the leader. 81.2% of the leading score, 31.19% of the leader’s cost.
  3. MiMo-V2-Flash (Feb 2026). Cheapest efficient configuration at 70% of the leader. 72.5% of the leading score, 7.65% of the leader’s cost.

Saved data (JSON) · Artificial Analysis

The staircase is the Pareto frontier: configurations for which no other option scores at least as well for less, or better for the same cost. March had efficient choices. None cleared our 80/20 threshold. A place on the staircase and a place in the challenger box are separate qualifications.

September: 29 challengers

In the saved 4 September snapshot, Claude Fable 5.1 at max effort held the highest score. Apply the same relative rule and 29 of 129 eligible configurations sit inside the challenger box.

Different reasoning efforts count as separate configurations. Each of these 29 challengers clears the same score floor and cost ceiling. The leader and the evaluation mix changed between the two snapshots, so this compares each market against its own leader. It does not measure the price decline for a fixed level of capability.

29 challengers beneath the crown, 4 September 202629 80/20 challengers among 129 eligible configurations. Shading marks the challenger box. The staircase is the Pareto frontier. The cost axis is logarithmic. Hover or tap a point for its model. With the chart focused, use the left and right arrow keys to inspect points.Score relative to the leader (%)0204060801000.1%1%10%100%1000%GPT-5.6 Luna (Non-reasoning). index: 26.8. index cost: $10.21.Llama 4 Scout. index: 10.3. index cost: $11.60.Granite 4.2 3B. index: 14.3. index cost: $12.45.GPT-5.6 Luna (low). index: 33.9. index cost: $13.94.GPT-5.6 Luna (medium). index: 38.9. index cost: $21.49.gpt-oss-120b (low). index: 14.9. index cost: $24.13.MiMo-V2.5. index: 38.0. index cost: $25.18.gpt-oss-20b (high). index: 15.2. index cost: $32.69.Gemma 4 31B (Non-reasoning). index: 22.3. index cost: $32.80.Granite 4.2 8B. index: 19.6. index cost: $33.49.Llama 4 Maverick. index: 14.5. index cost: $35.03.HyperNova 60B 2605 (high, based on gpt-oss-120b). index: 18.3. index cost: $37.50.NVIDIA Nemotron 3 Nano 30B A3B (Reasoning). index: 14.5. index cost: $39.36.Gemma 4 26B A4B (Reasoning). index: 26.1. index cost: $53.86.Granite 4.2 30B. index: 23.7. index cost: $55.58.GPT-5.6 Luna (high). index: 47.0. index cost: $55.86.Agnes 2.5 Pro Beta. index: 49.1. index cost: $56.11.Mistral Large 3. index: 15.9. index cost: $71.14.Nemotron 3.5 Lightning. index: 23.6. index cost: $72.11.Ling 3.0 Flash. index: 37.8. index cost: $72.81.Qwen3 Next 80B A3B (Reasoning). index: 16.9. index cost: $83.97.Mistral Small 4 (Reasoning). index: 19.7. index cost: $87.67.Ministral 3 3B. index: 7.1. index cost: $91.59.Hy3. index: 42.2. index cost: $93.20.Muse Glimmer (high). index: 35.1. index cost: $93.78.gpt-oss-120b (high). index: 24.1. index cost: $95.05.GPT-5.6 Luna (xhigh). index: 50.1. index cost: $96.36.Mercury 2. index: 21.9. index cost: $97.52.GPT-5.6 Terra (Non-reasoning). index: 34.6. index cost: $98.52.MiMo-V2.5-Pro. index: 42.9. index cost: $98.86.DeepSeek V4 Pro (high). index: 43.7. index cost: $108.77.Ministral 3 8B. index: 9.0. index cost: $114.06.Ministral 3 14B. index: 11.2. index cost: $119.20.GPT-5.6 Terra (low). index: 41.3. index cost: $132.83.GLM-5.3-Flash. index: 57.5. index cost: $138.02.Solar Pro 3. index: 14.5. index cost: $147.93.Gemini 3.5 Flash-Lite. index: 37.4. index cost: $153.39.Celeris-1. index: 12.4. index cost: $156.79.Qwen3.8-Flash-Next. index: 55.8. index cost: $158.13.Gemini 3.7 Flash (low). index: 50.9. index cost: $165.16.GPT-5.6 Luna (max). index: 52.3. index cost: $173.85.DeepSeek V4 Pro (max). index: 45.3. index cost: $178.12.GPT-5.6 Sol (Non-reasoning). index: 41.9. index cost: $181.24.Inkling Small. index: 41.2. index cost: $192.52.GPT-5.6 Terra (medium). index: 46.8. index cost: $196.62.Grok 4.3 (Non-reasoning). index: 25.0. index cost: $200.93.MiniMax-M3. index: 45.4. index cost: $204.82.Qwen3.5 122B A10B (Non-reasoning). index: 28.2. index cost: $221.35.Gemini 3.8 Flash (low). index: 51.7. index cost: $227.43.Qwen3.5 35B A3B (Non-reasoning). index: 24.3. index cost: $228.36.DeepSeek V4 Flash Vision (max). index: 51.5. index cost: $235.89.Qwen3.5 9B (Reasoning). index: 21.8. index cost: $240.43.Qwen3 Coder Next. index: 21.3. index cost: $250.18.Nemotron 3 Super 120B A12B (Reasoning). index: 25.7. index cost: $253.65.GPT-5.6 Sol (low). index: 50.7. index cost: $255.73.Grok 4.6 (low). index: 51.7. index cost: $260.34.Nex-N2-Pro. index: 41.7. index cost: $267.29.Qwen3.7 Plus. index: 39.4. index cost: $274.77.Gemini 3.7 Flash (medium). index: 53.4. index cost: $277.50.Kimi K3 (low). index: 48.3. index cost: $282.64.Trinity Large Thinking. index: 18.7. index cost: $289.19.Magistral Small 1.2. index: 11.5. index cost: $300.78.Qwen3.8 27B (low). index: 42.9. index cost: $302.62.Step 3.7 Flash. index: 30.9. index cost: $319.97.DeepSeek V4 Flash 0731 (max). index: 51.8. index cost: $323.26.Qwen3.6 27B (Non-reasoning). index: 31.3. index cost: $359.14.Solar Pro 4. index: 41.6. index cost: $380.17.Qwen3.8 27B (medium). index: 44.5. index cost: $398.33.GPT-5.6 Terra (high). index: 50.1. index cost: $405.59.Gemini 3.6 Flash (high). index: 51.6. index cost: $412.03.Claude Sonnet 5 (Non-reasoning, High Effort). index: 42.6. index cost: $415.75.Qwen3.8 27B (Non-reasoning). index: 34.7. index cost: $421.76.GPT-5.5 Instant (June 2026). index: 29.2. index cost: $424.21.GPT-5.6 Sol (medium). index: 55.6. index cost: $429.35.LongCat 2.0. index: 34.0. index cost: $441.77.Qwen3.5 122B A10B (Reasoning). index: 32.8. index cost: $443.97.Ring-2.6-1T. index: 31.7. index cost: $455.40.Gemini 3.8 Flash (medium). index: 56.6. index cost: $462.13.Nemotron 3 Ultra 550B A55B (Reasoning). index: 38.3. index cost: $470.33.Claude 4.5 Haiku (Reasoning). index: 29.9. index cost: $471.22.Gemini 3.7 Flash (high). index: 56.0. index cost: $484.73.Qwen3.6 35B A3B (Reasoning). index: 32.1. index cost: $505.33.Qwen3.5 397B A17B (Reasoning). index: 34.3. index cost: $527.22.Qwen3.6 35B A3B (Non-reasoning). index: 24.6. index cost: $529.96.Kimi K2.7 Code. index: 43.0. index cost: $544.58.Claude Opus 5 (low). index: 52.5. index cost: $555.25.GPT-6 Astra (low). index: 56.7. index cost: $574.92.GPT-5.6 Terra (xhigh). index: 52.8. index cost: $604.17.DeepSeek V4 Pro 0813 (max). index: 53.2. index cost: $616.95.Apodex 1.1. index: 44.0. index cost: $619.48.Grok 4.5 (high). index: 55.8. index cost: $628.65.Muse Spark 1.2 (xhigh). index: 56.8. index cost: $639.27.Qwen3.6 27B (Reasoning). index: 37.7. index cost: $667.18.Qwen3.8 27B (xhigh). index: 52.0. index cost: $693.67.Inkling (xhigh). index: 42.3. index cost: $698.28.GPT-5.6 Sol (high). index: 57.3. index cost: $699.68.Muse Spark 1.3 (xhigh). index: 60.8. index cost: $810.19.Gemini 3.1 Pro Preview. index: 47.7. index cost: $811.90.Gemini 3.8 Flash (high). index: 58.7. index cost: $825.83.Mistral Medium 3.5. index: 30.4. index cost: $843.20.GLM-5.2 (max). index: 52.6. index cost: $843.44.Grok 4.6 (medium). index: 59.0. index cost: $929.97.GPT-6 Astra (Non-reasoning). index: 55.3. index cost: $991.51.GPT-6 Astra (medium). index: 59.2. index cost: $1,032.97.Quasar 438B (max, based on GLM-5.2). index: 43.0. index cost: $1,047.71.Claude Fable 5.1 (low). index: 58.1. index cost: $1,086.44.GPT-5.6 Sol (xhigh). index: 59.0. index cost: $1,107.55.Claude Opus 5 (medium). index: 58.6. index cost: $1,116.24.Grok 4.6 (high). index: 60.9. index cost: $1,157.64.GLM-5.3 (max). index: 59.5. index cost: $1,238.50.Qwen3.8 2.4T A95B. index: 57.7. index cost: $1,395.05.GPT-5.6 Terra (max). index: 56.6. index cost: $1,409.48.Grok 4.6 (xhigh). index: 60.0. index cost: $1,428.64.GPT-6 Astra (high). index: 60.3. index cost: $1,433.33.Qwen3.8 Max. index: 58.1. index cost: $1,532.07.Claude Fable 5.1 (medium). index: 60.5. index cost: $1,572.41.Claude Opus 5 (high). index: 61.5. index cost: $1,974.35.GPT-6 Astra (xhigh). index: 61.0. index cost: $2,004.12.GPT-5.6 Sol (max). index: 60.9. index cost: $2,017.29.Claude Fable 5.1 (high). index: 62.5. index cost: $2,385.77.Kimi K3 (max). index: 59.7. index cost: $2,425.11.Claude Opus 5 (xhigh). index: 62.5. index cost: $2,909.30.GPT-6 Astra (max). index: 61.2. index cost: $3,020.82.Muse Spark 1.3 (max). index: 62.1. index cost: $3,049.64.Claude Opus 5 (max). index: 63.1. index cost: $3,836.05.Claude Sonnet 5 (max). index: 55.3. index cost: $4,010.51.Claude Fable 5.1 (xhigh). index: 64.8. index cost: $5,220.73.Claude Fable 5 (max). index: 62.1. index cost: $5,455.22.Claude Fable 5.1 (max). index: 65.7. index cost: $8,523.16.123Cost relative to the leader (%)

29 80/20 challengers / 129 eligible configurations. Shading marks the challenger box. The staircase is the Pareto frontier. The cost axis is logarithmic.

  1. Claude Fable 5.1 (max). Leading score. 65.7 index, $8,523.16 index cost.
  2. GLM-5.3-Flash. Cheapest efficient configuration at 80% of the leader. 87.5% of the leading score, 1.62% of the leader’s cost.
  3. GPT-5.6 Luna (high). Cheapest efficient configuration at 70% of the leader. 71.5% of the leading score, 0.66% of the leader’s cost.

Saved data (JSON) · Artificial Analysis

Three points on the staircase show the whole argument. Claude Fable 5.1 at max effort is the frontier. GLM-5.3-Flash is the challenger: 87.5% of the leading score for 1.6% of the leader’s index cost. GPT-5.6 Luna at high effort sits at the bottom of the staircase, the cheapest configuration that keeps at least 70% of the leading score, at 0.66% of the cost. All three are Pareto-efficient. Which one is the Pareto model depends on where the workload sets its bar.

GLM-5.3-Flash makes the challenger concrete, and the team behind it has been building toward this for years. Z.ai, formerly Zhipu AI, grew out of Tsinghua University. Nathan Lambert’s history of GLM at Interconnects traces the company’s founding in 2019, the first GLM research in 2021, and ChatGLM in 2023. His reading of the current generation is that it is a post-training story, not a distillation story: the same base model, with more environments and more compute spent training on them. This is a sustained research program showing up in the price breakers’ game.

GLM-5.3-Flash has 320 billion parameters, with 18 billion active per token, and combines sparse and linear attention. It takes text, images, and video as input. Those are specific architectural choices behind a model with an unusually small benchmark bill. They earn it a place in the challenger box. Whether it becomes my Pareto model depends on the work it can deliver.

The task sets the budget

Fable is the reliable engineer in my setup. When designing a new system, working through an unfamiliar codebase, or deciding what to build, I want to maximize the intelligence I can get. A better model might find the simpler architecture, catch a broken assumption, or save me a day on the wrong problem. That is specifically what the frontier premium gets me. The opportunity cost of building a worse system far outweighs the cost of the maximum intelligence that might prevent it.

On the contrary, if the model is part of a system, the calculus changes. Inside an engineered system, the model works with constraints and specific instructions. The model receives a contract to follow, and novelty is a defect. Inside such a system, evaluations are run on the model at scale to reach a probabilistic conclusion that the model is suitably instructed, with errors prevented through retry loops, for example. Once a cheaper model meets that contract reliably, the costly flagship has little to no justified presence.

My default is Frontier for discovery, Pareto for iterated delivery. Pay for judgement while the problem is open. Once the problem is defined and a cheaper model is proven, settle for the cheapest option that meets the bar.

What comes after the price breakers?

Over the next year, I expect frontier models to take on longer design and engineering jobs from an incomplete brief, and to become more able to correct their own course and explore decision branches. For Pareto models, I expect greater speed and a lower price per task.

These are roles in a system. The same lab, and sometimes the same model at different reasoning settings, may serve both.

Recent pricing changes show why neither class will have one simple price:

  • Anthropic estimates that Fable 5.1 will cost 25% less than Fable 5 on typical token-billed workloads because of cheaper cache reads, with savings up to approximately 45% for highly agentic work. Fable 5.1 announcement.
  • Claude Fable 5.1 through Anthropic's Message Batches API costs $5 per million input tokens and $25 per million output tokens. Exactly 50% of its standard $10 and $50 rates. Requests run asynchronously and can take up to 24 hours. Fable 5.1 pricing, Message Batches API.
  • DeepSeek announced peak and off-peak billing effective 16 August 2026, with off-peak rates half the peak rates. DeepSeek pricing update.
  • OpenAI lists Sol’s promotional pricing as available at least through 21 November 2026. Google Cloud lists introductory standard global Gemini 3.7 Flash rates of $0.75 input and $3.75 output per million tokens through December, with $1.50 and $7.50 scheduled for 1 January 2027. OpenAI pricing, Google Cloud pricing.
  • GPT-6 Astra prices context length and urgency separately. Standard rates are $10 per million input tokens and $50 per million output tokens. A request with more than 272,000 input tokens is billed at $20 and $75 for the whole request. Batch and Flex cost half the applicable standard rates. Fast mode costs twice them. Astra pricing.

I expect model selection to become more tightly tied to the type of task the model is applied to. Reusable context, a flexible delivery window, or the end of a promotion can change the cost of a task. A stronger frontier model can earn a larger bill if it saves equivalent human work.

Model selection can thus be consolidated into creation versus iteration. A Frontier model is the creator, the Pareto model the iterator.

Sources and measurement

The figures use the 6 March leaderboard capture and the Artificial Analysis leaderboard retrieved on 4 September 2026. These are fixed observations, not a live feed.

The chart dataset, March observations, September observations, and pricing source notes are available as JSON. The chart dataset records source URLs, file hashes, and inclusion counts.

The definition and batch-pricing notes record the additional sources checked through Exa on 6 September 2026. The publication source notes cover the Fable 5 redeployment, the GLM history and architecture, and the Astra pricing. The workload definitions are our editorial choices. The batch rates are Anthropic's published prices.

Both figures are also available as images, with the three highlighted points labelled: March as PNG or SVG, September as PNG or SVG. The highlights follow one rule on both dates: the leader, then the cheapest efficient configuration keeping 80% of its score, then the cheapest keeping 70%.

Each point is a non-deprecated configuration with a non-estimated score and a positive recorded cost for running that snapshot's full index. We exclude configurations without those measurements. The reference is the highest-scoring eligible configuration, with lower cost breaking ties. Both plots use the same relative axes.

The score percentage is a fraction of a benchmark index, not a fraction of intelligence. Cost includes the evaluation's token usage at the recorded prices. It is not a quote for a production workload. See the Artificial Analysis methodology.

Two September costs were calculated from Artificial Analysis token counts and external prices: Muse Spark 1.3 (max) from Meta, and Nex-N2-Pro from OpenRouter. Their saved records identify that calculation. Motif 3 and K2 Horizon 375B remain excluded because the saved work did not establish a public price.

Research