I keep researching this for my own work, so I turned it into a reference. One big table, everything in it: prices, benchmarks, licenses, where to access each model. Scan it, pick what you need, go build.

Quick mental model before the table: Agent = Model + Harness. The model is the intelligence, the harness (Claude Code, Cursor, Cline, Aider...) turns it into something that navigates your repo and fixes its own mistakes. Pick both.

The table

ModelCreatorLicensePrice in/out per 1MSWE-BenchDirect accessAlso on
Claude Fable 5AnthropicClosed$10 / $5080.3% Proplatform.claude.comClaude Code, claude.ai
Claude Opus 4.8AnthropicClosed$5 / $2569.2% Pro, 88.6% Verifiedplatform.claude.comClaude Code, Cursor, OpenRouter, Bedrock, Vertex
Claude Sonnet 5AnthropicClosed~$3 / $15new, near Opus 4.8platform.claude.comClaude Code, Cursor
Claude Sonnet 4.6AnthropicClosed$3 / $15strongplatform.claude.comCursor, Windsurf, Cline, OpenRouter
Claude Haiku 4.5AnthropicClosed$1 / $573.3% Verifiedplatform.claude.comOpenRouter
GPT-5.5OpenAIClosed$5 / $3058.6% Pro, 88.7% Verifiedplatform.openai.comCodex CLI, Copilot, Cursor, Azure
GPT-5.3-CodexOpenAIClosed$1.75 / $14coding-tunedplatform.openai.comCodex CLI
Gemini 3.1 ProGoogleClosed$2 / $12 ($4 / $18 above 200K)54.2% Pro, 80.6% Verifiedaistudio.google.comGemini CLI, Vertex, Cursor, OpenRouter
Grok 4.xxAIClosed~$2.60 / $7.80Tier Bx.aiCursor, OpenRouter
DeepSeek V4-ProDeepSeekMIT$1.74 / $3.4880.6% Verifiedplatform.deepseek.comOpenRouter, Together, Morph, HuggingFace
DeepSeek V4-FlashDeepSeekMIT$0.14 / $0.2879% Verifiedplatform.deepseek.comOpenRouter (:free tier), Morph, Ollama
DeepSeek V3.2DeepSeekMIT~$0.23 / ~$1best value classicplatform.deepseek.comOpenRouter, Ollama
GLM 5.2Z.aiOpen$1.40 / $4.40 (~$0.45 / $3.31 avg on OR)62.1% Pro (vendor)z.aiOpenRouter, HuggingFace, GLM Coding Plan
GLM-4.7-FlashZ.aiOpenfreedecentz.aiOpenRouter, Ollama
Kimi K2.6MoonshotOpen$0.95 / $4.0080.2% Verified, 58.6% Proplatform.moonshot.aiOpenRouter, Together, Groq, HuggingFace
Kimi K2.7 CodeMoonshotApache 2.0~K2.6 rates, 30% fewer thinking tokenscoding-tunedplatform.moonshot.aiOpenRouter, HuggingFace
MiniMax M3MiniMaxOpen$0.60 / $2.40 (~$0.10 / $1.21 avg on OR)80.5% Verifiedminimax.ioOpenRouter, Atlas Cloud
MiniMax M2.7MiniMaxOpen$0.30 / $1.2056.2% Prominimax.ioTogether, Atlas Cloud
Qwen3.7-MaxAlibabaClosed APIvia Alibaba Cloud80%+ Verifiedalibabacloud.comOpenRouter
Qwen 3.6-27BAlibabaApache 2.0free local (22GB VRAM)77.2% VerifiedOllama, HuggingFaceTogether, OpenRouter
Qwen 3.6 PlusAlibabaOpen$0.50 / $3.0061.6% Terminal-BenchAlibaba CloudTogether, OpenRouter
Qwen3 CoderAlibabaApache 2.0free on OpenRouterbest free coding modelOpenRouter :freeOllama
Qwen 2.5 Coder 32BAlibabaApache 2.0free local (18GB+ VRAM)best local classicOllamaLM Studio, Continue.dev
Nemotron 3 UltraNVIDIAOpenMDW~$0.42 / $2.61 (OR avg)#2 open on AA indexNVIDIA NIMOpenRouter (:free route), HuggingFace
North Mini CodeCohereApache 2.0free local33.4 AA Coding IndexHuggingFaceOllama, vLLM
Codestral 22BMistralOpenfree local (12GB VRAM)best autocomplete localOllamaMistral API, OpenRouter
Devstral Small 24BMistralOpenfree local (16GB VRAM)near-frontier localOllamaMistral API
Llama 4 ScoutMetaOpenfree local, :free on ORsolid, 10M contextHuggingFaceOllama, Groq, Together
Poolside Laguna M.1PoolsideClosedfree tier#1 on Kilo usageKilo CodePoolside platform
openPangu 2.0HuaweiOpenfree localcompetitiveHuggingFacevLLM self-host

How to read it:

  • Prices are first-party rates as of late June 2026. OR = OpenRouter weighted average, noted where it differs a lot from direct.
  • OpenRouter adds ~5.5% credit fee ($0.80 minimum) and 5% BYOK fee above 1M requests/month. The per-token rates themselves match provider list prices.
  • Benchmarks mix vendor-reported and independent numbers. Compare within the same column only, and test on your own code before committing.
  • Fable 5 caveat: a US export-control directive suspended access on June 12, expected back for US users around July 1. If you are outside the US (like me), verify before building on it.

The agents (harnesses)

AgentTypeModelsLink
Claude CodeFirst-party CLI/appClaude familyclaude.com/claude-code
Codex CLIFirst-party CLIGPT familyopenai.com/codex
Gemini CLIFirst-party CLIGeminigithub.com/google-gemini/gemini-cli
Copilot Agent ModeIDE-nativeGPT familygithub.com/features/copilot
CursorAI IDEBYOM (Claude, GPT, Gemini...)cursor.com
WindsurfAI IDEMultiplewindsurf.com
ClineVS Code ext, BYOMAnything via OpenRouter/Ollamacline.bot
RooCodeVS Code ext, BYOMAnything, strong on large multi-file workroocode.com
AiderTerminal, BYOMAnything, git-nativeaider.chat
Kilo CodeBYOM, free tiersAnything, live usage leaderboardkilo.ai
Continue.devVS Code ext, localOllama modelscontinue.dev

Providers cheat sheet

ProviderWhat it isWhen to use
First-party APIsDirect from each creatorCheapest per token on one provider
OpenRouter315+ models, one keyMulti-model access, budget +5-7% overhead
Together AI / FireworksNeutral open-model hostsOpen models with fine-tuning, dedicated deploys
GroqFast inferenceSpeed on open models
Morphbf16, no quantizationOpen-model fidelity for codegen (most hosts quantize to fp8 and lose quality)
Bedrock / Vertex / AzureCloud resellers10-20% more per token, but compliance and one cloud bill
OllamaLocal runnerFree, private, offline
SubscriptionsClaude Pro/Max, Codex Plus, GLM Coding Plan, OpenCode Go ($10/mo)Can beat API billing depending on your usage profile

Three numbers to remember

The 10x cliff. Five models score between 80.2% and 80.6% on SWE-bench Verified (DeepSeek V4 Pro, Gemini 3.1 Pro, MiniMax M3, Qwen3.7 Max, Kimi K2.6), with output prices from $2.40 to $12 per million. The next 8 points up (GPT-5.5, Opus 4.8) cost $25 to $30. The last 8% of quality costs 10x. Know if your work lives in that gap.

The 21.7 point gap. Claude Fable 5 scores 80.3% on SWE-Bench Pro vs 58.6% for GPT-5.5. Frontier benchmarks usually move in single digits. This one did not.

1/36th. DeepSeek V4 Flash delivers about 82% of Opus's SWE-Bench Pro score at roughly 1/36th the input price. It does not replace Opus. It means Opus should stop doing DeepSeek's job.

My routing shortcut

  • Cheap and low-risk work: DeepSeek V4 Flash
  • Default coding workhorse: Kimi K2.6 or Sonnet-class
  • Hard architecture, repeated failures: frontier (Opus 4.8, GPT-5.5, Fable 5 if you can get it)
  • Private or offline: Qwen local via Ollama
  • Always: measure cost per successful task, not cost per token. A cheap model that causes rework is expensive.

This space moves monthly. I will update this table when the next shakeup lands. Building with agents and want to compare notes? I am @itseduvieira pretty much everywhere.