<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Signal Room</title><link>https://projectious-work.github.io/ai-market-research/</link><description>Recent content on Signal Room</description><generator>Hugo</generator><language>en</language><atom:link href="https://projectious-work.github.io/ai-market-research/index.xml" rel="self" type="application/rss+xml"/><item><title>01 Market &amp; Economics</title><link>https://projectious-work.github.io/ai-market-research/report/01-market-economics/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/report/01-market-economics/</guid><description>&lt;div class="sr-report-scope"&gt;
&lt;script id="market-data" type="application/json"&gt;{"meta":{"generated_at":"2026-08-07T00:00:00Z","reference_default":"fable-5","report_metrics_file":"data/report-metrics.json"},"executive_summary":{"models":["**Claude Opus 5 is now generally available.** Anthropic's new `claude-opus-5` targets complex agentic coding and enterprise work with a 1M-token context window, 128K maximum output, adaptive thinking by default, and unchanged $5/$25 per MTok base pricing. Independent benchmark values are not yet recorded here.","**Google reset the Flash price-performance curve on July 21.** Gemini 3.6 Flash is GA at $1.50/$7.50 per million tokens with stronger coding and agentic results than 3.5 Flash, while Gemini 3.5 Flash-Lite reaches roughly 490 output tok/s at $0.30/$2.50 for high-volume subagents and extraction.","**Kimi K3 resets the open-weight frontier.** Moonshot reports a 2.8T sparse MoE, native vision, 1M context, and 99% of Fable 5 across 14 overlapping launch-suite evaluations; weights are promised by July 27.","**GPT-5.6 Luna and Terra became materially cheaper on July 30.** Luna fell 80% to $0.20/$1.20 per MTok and Terra 20% to $2/$12; their paid Codex and ChatGPT Work usage also consumes fewer credits, while subscription prices and quota budgets did not change.","**GPT-5.6 now spans Sol, Terra and Luna**, giving leaders a deliberate capability, balanced, and high-throughput ladder under one family. Sol API Fast mode replaces Priority Processing: OpenAI claims up to 2.5× Standard speed at 2× price, with unchanged intelligence.","**Luna is expanding beyond API routing.** OpenAI says it becomes the default for Free and Go users, with unlimited text chats and a higher-reasoning Think option rolling out subject to abuse guardrails. This is a ChatGPT product update, not an API capability change.","**Keep benchmark tables separate by harness.** OpenAI's GPT-5.6 launch results and Scale's public SWE-bench Pro leaderboard use different model versions and evaluation setups; use each as evidence, not as one directly rankable series.","**GPT-5.5 remains active in Codex and the API.** It is retained as a compatibility and portfolio option rather than being hidden by the newer 5.6 family.","**Claude Haiku 4.5 remains Anthropic’s latest verified Haiku.** No official Haiku 5 listing was found in Anthropic’s current model catalog."],"harnesses":["**Separate model capability from harness capability.** Tool execution, context management, isolation, and observability can dominate real workflow outcomes.","**Subscription access is not a production routing contract.** Validate API, credit-pool, and third-party harness policies before standardizing an operating model.","**Maintain at least one portable fallback path** across providers for high-value workflows and operational incidents.","**Gemini Managed Agents now add background execution, remote MCP, custom functions and credential refresh.** That makes Google a more credible managed-agent control plane, but it does not substitute for workload-specific evaluation."],"self_hosting":["**Open weights are now a strategic option, not only a cost play.** Kimi K3, DeepSeek, Qwen, Nemotron, Mistral, and Llama cover different sovereignty and specialization needs.","**Do not compare self-hosting at zero token cost.** Include accelerator rental or depreciation, power, utilization, serving staff, and measured throughput.","**Pilot against a defined workload and hardware envelope** before treating advertised context or parameter scale as deployable capacity."],"strategy":["**Run a portfolio, not a winner-takes-all model standard.** Reserve frontier reasoning for high-value decisions and route routine work to measured fast or efficient tiers.","**Instrument quality, latency, retries, and total workflow cost together.** Token price alone is not an operating metric.","**Review the portfolio quarterly and after major releases**, with explicit retirement, security, and fallback criteria."]},"headline_stats":[{"id":"frontier_count","label":"Models tracked","value":35,"unit":"","delta":"Current, fast and retained compatibility models","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_weight","label":"Open-weight","color":"#16866f"}],"history":[{"date":"2026-05-18","proprietary":20,"open_weight":3},{"date":"2026-07-22","proprietary":26,"open_weight":8},{"date":"2026-07-25","proprietary":27,"open_weight":8}],"note":"Release-tag roster snapshots; announced open-weight releases are counted with open-weight models."}},{"id":"best_open_pct","label":"Best open-weight vs selected reference","value":99,"unit":"%","delta":"Kimi K3; geometric mean across 14 overlapping vendor evals","delta_dir":"up","signal_id":"open_weight_quality"},{"id":"cheapest_frontier_api","label":"Lowest frontier blended API price","value":0.9,"unit":"$/Mtok","delta":"GPT-5.6 Luna; 30% input / 70% output blend","delta_dir":"down","signal_id":"blended_frontier_price"},{"id":"fastest_task_rate","label":"Fastest task-rate index","value":"3.2×","unit":"","delta":"Fable 5 = 1.0×; planning index, not tokens/second","delta_dir":"up","signal_id":"task_speed"},{"id":"max_context","label":"Largest usable context","value":"10M","unit":"tok","delta":"Llama 4 Scout; deployment constraints still apply","delta_dir":"up","trend":[{"date":"2026-01-01","value":1},{"date":"2026-04-01","value":1},{"date":"2026-07-01","value":10}]},{"id":"harness_count","label":"Harnesses tracked","value":16,"unit":"","delta":"Commercial and open agent environments","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_source","label":"Open source","color":"#3d78c5"}],"history":[{"date":"2026-05-18","proprietary":3,"open_source":11},{"date":"2026-07-22","proprietary":6,"open_source":10}],"note":"Release-tag roster snapshots classified from each harness license."}},{"id":"documented_output_speed","label":"Fastest documented API output","value":490,"unit":"tok/s","delta":"Gemini 3.5 Flash-Lite; Artificial Analysis first-party API measurement","delta_dir":"up","signal_id":"output_throughput"},{"id":"quality_coverage","label":"Models with quality evidence","value":"31/35","unit":"","delta":"Unknown remains unknown; no zero-value substitution","delta_dir":"neutral","trend":[{"date":"2026-01-01","value":18},{"date":"2026-04-01","value":24},{"date":"2026-07-01","value":31}]}],"trends":{"best_open_vs_opus":{"label":"Best open-weight % vs Opus 4.7","history":[{"date":"2025-11-01","value":68},{"date":"2025-12-01","value":72},{"date":"2026-01-01","value":78},{"date":"2026-02-01","value":80},{"date":"2026-03-01","value":83},{"date":"2026-04-01","value":86},{"date":"2026-05-01","value":88},{"date":"2026-06-01","value":90}]},"median_frontier_output_price":{"label":"Median frontier $/Mtok output","history":[{"date":"2025-11-01","value":18},{"date":"2025-12-01","value":17},{"date":"2026-01-01","value":15.5},{"date":"2026-02-01","value":14.5},{"date":"2026-03-01","value":13},{"date":"2026-04-01","value":12.5},{"date":"2026-05-01","value":12},{"date":"2026-06-01","value":11}]},"models_released_per_month":{"label":"Notable model releases per month","history":[{"date":"2025-11-01","value":3},{"date":"2025-12-01","value":4},{"date":"2026-01-01","value":5},{"date":"2026-02-01","value":6},{"date":"2026-03-01","value":4},{"date":"2026-04-01","value":7},{"date":"2026-05-01","value":5},{"date":"2026-06-01","value":8}]}},"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the GPT-5.6 pricing history from OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. The report now records lower paid Codex and ChatGPT Work credit consumption, unchanged subscription prices and quota budgets, Sol API Fast mode (up to 2.5× Standard speed at 2× price), and the August 6 ChatGPT Luna access expansion."},{"date":"2026-08-01","tag":"pricing","text":"Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to $0.20/$1.20 from OpenAI's live model pages. Recomputed the tracked 30/70 workload blends and effort-burn ratios. Superseded on August 7 with OpenAI's published July 30 effective date and full price-change details."},{"date":"2026-08-01","tag":"model","text":"Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output, thinking/non-thinking modes, $0.14/$0.28 per-million pricing, $0.0028 cached input, and vendor-reported agent results including Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2."},{"date":"2026-07-28","tag":"model","text":"Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B active) open-weight (Apache 2.0) multimodal MoE with 1M-token context, released 2026-07-15. No first-party per-token API pricing is published (monetized via the Tinker fine-tuning platform); benchmark scores are published on the model card but not yet normalized into this roster's comparable set, so quality/speed/cost fields are left unknown rather than estimated."},{"date":"2026-07-25","tag":"model","text":"Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive thinking by default, $5/$25 per MTok base pricing, and official cloud-platform availability. Added independent Artificial Analysis evidence (61 Intelligence Index at max effort; 52.3 output tok/s) with effort-specific caveats. Gemini 3.5 Flash Cyber remains limited to CodeMender government and trusted-partner pilots; GPT-Live and Muse Spark 1.1 remain non-API products, so none were added to the API roster. Claude Opus 4.7 Fast Mode was removed July 24."},{"date":"2026-07-24","tag":"benchmark","text":"Refreshed current-source evidence: added OpenAI's GPT-5.6 launch table, Google's Managed Agents update, and Scale's public SWE-bench Pro leaderboard. Clarified that vendor launch tables and the public leaderboard are not directly comparable because their model versions and harnesses differ."},{"date":"2026-07-22","tag":"fix","text":"Made the selected reference propagate through open-weight quality headlines, market-signal history, comparison headings, model analytics, self-hosting quality and capability market position. Replaced the model and harness inventory mini-lines with stacked proprietary/open category areas based on release-tag roster snapshots."},{"date":"2026-07-22","tag":"model","text":"Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M context, 64K output, current API pricing, Artificial Analysis intelligence and throughput measurements, and Google’s published coding and agentic benchmarks."},{"date":"2026-07-22","tag":"data","text":"Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to governments and trusted CodeMender partners; Gemini Omni Flash and Nano Banana 2 Lite remain specialized media models rather than general-purpose roster entries."},{"date":"2026-07-21","tag":"model","text":"Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint: fully open Apache 2.0 weights and data, 4K context, 1.7T-token distillation, 1,811 languages, and official BF16, FP8, NVFP4A16, INT3, INT4 and INT6 variants."},{"date":"2026-07-21","tag":"harness","text":"Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI 0.144.6, Gemini CLI 0.51.0, OpenCode 1.18.4, Cline 4.0.10, and Goose 1.43.0; updated repository star snapshots and notable release capabilities."},{"date":"2026-07-21","tag":"model","text":"Added Moonshot's official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached input, and $15/M output; the 30/70 workload blend is $11.40/M before reasoning-effort effects."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark registers and selectable reference configurations. Added documented quality, speed, cost and capability composites in data/report-metrics.json."},{"date":"2026-07-18","tag":"routing","text":"Updated the action queue and recommended routing: Fable for hardest retained-data workloads, Terra for default engineering, Luna for high-volume subagents, with explicit escalation rules."},{"date":"2026-06-06","tag":"fix","text":"Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K context, 1000+ tok/s on Cerebras) and Hermes Agent v0.16.0 (Nous Research, MIT, self-hosted multi-platform agent) — both were incorrectly removed in v1.5.0 sweep."},{"date":"2026-06-06","tag":"policy","text":"Dashboard market sweep v1.5.0: real-world re-grounding. Replaced fictional Mythos/GPT-5.5-Cyber rows with verified models; added Nvidia Nemotron coalition, Kimi K2.6, GLM-5, Cohere Command A+, SubQ 1M-Preview."},{"date":"2026-06-04","tag":"model","text":"Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context, NVIDIA Open Model License) at Computex — first frontier-scale open model from Nvidia."},{"date":"2026-06-04","tag":"model","text":"Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral, Perplexity, Reflection AI, Sarvam, Thinking Machines Lab as inaugural members."},{"date":"2026-06-01","tag":"model","text":"Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model."}],"actions":["P0 · Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","P0 · Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","P0 · Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","P1 · Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","P1 · Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","P2 · Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics."],"models":[{"id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","tier":"frontier","released":"2026-06-09","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":80,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii":59.9,"coding_agent_index":77.2,"deep_swe":69.7,"terminal_bench":83.1,"agents_last_exam":40.5,"api_in":10,"api_out":50,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":52.3,"subscription":"Claude API and supported cloud platforms","notes":"Default reference for v2. Anthropic describes Fable 5 as its most capable widely released model. Adaptive thinking is always on. Comparative latency is slower. Retention is 30 days and zero-data-retention is not available.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"status":"stable","speed_class":"deliberate","speed_evidence":"vendor-qualitative","capability_levels":{"coding":49,"reasoning":50,"knowledge":48,"comms":48,"multimodal":45,"agentic":50}},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":64.6,"swe_verified":null,"aaii":58.9,"coding_agent_index":80,"deep_swe":72.7,"terminal_bench":88.8,"agents_last_exam":52.7,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Flagship GPT-5.6 tier. API Fast mode replaced Priority Processing on July 30: OpenAI claims up to 2.5× Standard speed at 2× Standard price with no intelligence change. Speed is otherwise stored as an end-to-end task index in report-metrics.json, not tok/s.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":49,"knowledge":48,"comms":48,"multimodal":48,"agentic":49}},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":63.4,"swe_verified":null,"aaii":55,"coding_agent_index":77.4,"deep_swe":69.6,"terminal_bench":87.4,"agents_last_exam":50.4,"api_in":2,"api_out":12,"api_cache_hit":0.2,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Balanced GPT-5.6 tier and recommended default engineering route. On July 30, API list pricing fell 20% from $2.50/$15 to $2/$12 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":48,"reasoning":46,"knowledge":46,"comms":46,"multimodal":45,"agentic":48}},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","tier":"fast","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":62.7,"swe_verified":null,"aaii":51.2,"coding_agent_index":74.6,"deep_swe":67.2,"terminal_bench":84.7,"agents_last_exam":50.3,"api_in":0.2,"api_out":1.2,"api_cache_hit":0.02,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Fastest, most affordable GPT-5.6 tier; recommended for high-volume subagents with verification and escalation. On July 30, API list pricing fell 80% from $1/$6 to $0.20/$1.20 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change. ChatGPT is also rolling Luna out as the Free and Go default, with unlimited text chats and a Think option; that product change does not alter API routing.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":46,"reasoning":44,"knowledge":43,"comms":45,"multimodal":43,"agentic":46}},{"id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","tier":"frontier","released":"2026-07-16","license":"open weights announced; terms pending weight release","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":null,"tok_per_sec":null,"subscription":"Kimi API, Kimi Code, Kimi Work","reasoning_capable":true,"effort_default":"max","effort_levels":["max"],"notes":"2.8T sparse MoE with 16/896 experts active, native vision and always-on thinking. API pricing is $3/M uncached input, $0.30/M cached input and $15/M output. Full weights promised by July 27, 2026.","quality_vs_fable":99.2,"quality_evidence":"Geometric mean across 14 overlapping values in Moonshot’s launch comparison; vendor-reported, max/xhigh settings.","capability_levels":{"coding":47,"reasoning":46,"knowledge":47,"comms":46,"multimodal":47,"agentic":48}},{"id":"opus-5","name":"Claude Opus 5","provider":"Anthropic","tier":"frontier","released":"2026-07-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"Claude API, Bedrock, Google Cloud, Microsoft Foundry","notes":"Current Anthropic Opus generation. 1M context and 128K maximum output. Adaptive thinking is enabled by default; effort defaults to high. Artificial Analysis reports a 61 Intelligence Index and 52.3 output tok/s at max effort; these figures are effort-specific. Research-preview Fast Mode is Claude API-only at $10/$50 per MTok and claims up to 2.5x higher output throughput.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"speed_class":"balanced","speed_evidence":"vendor-qualitative","cite":[86,87]},{"id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","tier":"frontier","released":"2026-05-28","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":69.2,"swe_verified":88.6,"livecodebench":82,"aime":90,"tau2":86,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Max 20× $200/mo · Max 5× $100/mo","notes":"Previous Opus generation, retained as an active compatibility option. SWE-V 88.6%, SWE-Pro 69.2%, AAII 61.4. Fast Mode reduced 3× to $10/$50 (was $30/$150 on 4.7). 1M context standard.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":50,"comms":50,"multimodal":34,"agentic":50}},{"id":"opus-4.7","name":"Claude Opus 4.7","provider":"Anthropic","tier":"frontier","released":"2026-04-16","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":64.3,"swe_verified":87.6,"livecodebench":79,"aime":88,"tau2":84,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":50,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Now legacy as of Opus 4.8 release May 28. Pricing unchanged. SWE-Verified 87.6%, SWE-Pro 64.3%. Fast Mode was removed July 24, 2026; standard-speed API access remains active.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":49,"comms":50,"multimodal":32,"agentic":50}},{"id":"sonnet-4.6","name":"Claude Sonnet 4.6","provider":"Anthropic","tier":"frontier","released":"2026-02-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":44,"swe_verified":77,"livecodebench":76,"aime":85,"tau2":81,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":50,"tok_per_sec":80,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Best code style/intent understanding. With cache+batch: $0.30/$7.50 effective.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","max"],"cache_discount":0.1,"capability_levels":{"coding":41,"reasoning":40,"knowledge":39,"comms":42,"multimodal":31,"agentic":41}},{"id":"haiku-4.5","name":"Claude Haiku 4.5","provider":"Anthropic","tier":"fast","released":"2025-10-15","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":28,"swe_verified":62,"livecodebench":58,"aime":70,"tau2":65,"api_in":1,"api_out":5,"api_cache_hit":0.1,"batch_discount":50,"tok_per_sec":110,"subscription":"Available in all tiers","notes":"5× cheaper than Sonnet. Triage/classification champion.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.1,"capability_levels":{"coding":28,"reasoning":18,"knowledge":29,"comms":30,"multimodal":18,"agentic":28}},{"id":"gpt-5.5","name":"GPT-5.5","provider":"OpenAI","tier":"frontier","released":"2026-04-23","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M (272K threshold)","swe_pro":58.6,"swe_verified":88.7,"livecodebench":84,"aime":92,"tau2":82,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Pro $200 · Pro Lite $100 · Plus $20","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"OpenAI flagship; default in ChatGPT (Instant variant since May 5, 2026). 1M context with 2× input/1.5× output surcharge above 272K. Reasoning tokens billed as output.","cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":41,"knowledge":49,"comms":40,"multimodal":41,"agentic":38}},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","provider":"OpenAI","tier":"frontier","released":"2026-04-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":60,"swe_verified":90,"livecodebench":86,"aime":94,"tau2":84,"api_in":30,"api_out":180,"api_cache_hit":3,"batch_discount":50,"tok_per_sec":35,"subscription":"Pro $200 only","reasoning_capable":true,"effort_default":"high","effort_levels":["medium","high","xhigh"],"notes":"Highest-stakes reasoning tier. $30/$180. Available in ChatGPT Pro $200 and as API. 6× cost of base 5.5.","cache_discount":0.25,"capability_levels":{"coding":47,"reasoning":48,"knowledge":50,"comms":42,"multimodal":42,"agentic":42}},{"id":"gpt-5.4","name":"GPT-5.4","provider":"OpenAI","tier":"frontier","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":57.7,"swe_verified":81,"livecodebench":81,"aime":90,"tau2":80,"api_in":2.5,"api_out":15,"api_cache_hit":0.25,"batch_discount":50,"tok_per_sec":60,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Best quality-per-credit. Held SWE-Pro lead Feb-April. Default for everyday coding.","cache_discount":0.1,"capability_levels":{"coding":40,"reasoning":40,"knowledge":40,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-mini","name":"GPT-5.4 Mini","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":38,"swe_verified":73,"livecodebench":71,"aime":78,"tau2":70,"api_in":0.4,"api_out":1.6,"api_cache_hit":0.04,"batch_discount":50,"tok_per_sec":130,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"94% of GPT-5.4 coding at 6× less. Best subagent. ~1/20 burn vs GPT-5.5. Collapses at 64K+ context.","cache_discount":0.1,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-nano","name":"GPT-5.4 Nano","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":25,"swe_verified":60,"livecodebench":58,"aime":65,"tau2":55,"api_in":0.1,"api_out":0.4,"api_cache_hit":0.01,"batch_discount":50,"tok_per_sec":200,"subscription":"API only","reasoning_capable":true,"effort_default":"low","effort_levels":["minimal","low","medium"],"notes":"Smallest reasoning model. API-only. For embeddable/edge inference at near-zero cost.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":20,"comms":20,"multimodal":20,"agentic":20}},{"id":"gpt-5.3-codex","name":"GPT-5.3-Codex","provider":"OpenAI","tier":"frontier","released":"2026-01-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":56.8,"swe_verified":85,"livecodebench":82,"aime":87,"tau2":78,"api_in":1.5,"api_out":10,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":70,"subscription":"All paid tiers · Code Review uses this","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Coding-specialised. ~⅓ burn vs GPT-5.5 for ~2pts less SWE-Pro. Often best $/quality for execution turns.","cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":28,"knowledge":28,"comms":27,"multimodal":8,"agentic":38}},{"id":"gpt-5.3-codex-spark","name":"GPT-5.3-Codex-Spark","provider":"OpenAI","tier":"fast","released":"2026-02-12","license":"proprietary","jurisdiction":"US","context":128000,"context_label":"128K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":1000,"subscription":"ChatGPT Pro — research preview only","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"Smaller, latency-first sibling of GPT-5.3-Codex. Released Feb 12, 2026. 1000+ tok/s on Cerebras hardware. ChatGPT Pro research preview only — not in API at launch; separate preview rate-limit pool (no standard credit burn). Text-only. Target use: real-time micro-edits, live pair-programming in Codex app/CLI/VS Code.","cache_discount":null,"capability_levels":{"coding":38,"reasoning":22,"knowledge":22,"comms":25,"multimodal":0,"agentic":28}},{"id":"gpt-5.2","name":"GPT-5.2","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":52,"swe_verified":78,"livecodebench":76,"aime":84,"tau2":75,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Available but not recommended","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Codex picker keeps it for 'long-running agents' (specifically tuned for autonomy). Otherwise eclipsed by 5.3-Codex.","cache_discount":0.1},{"id":"gpt-5.2-codex","name":"GPT-5.2-Codex","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":50,"swe_verified":76,"livecodebench":75,"aime":82,"tau2":73,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Legacy","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Predecessor to 5.3-Codex. Legacy.","cache_discount":0.1},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","provider":"OpenAI","tier":"fast","released":"2026-05-05","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":35,"swe_verified":76,"livecodebench":72,"aime":81,"tau2":70,"api_in":1.5,"api_out":6,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":200,"subscription":"Default ChatGPT model · API chat-latest","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"ChatGPT default since May 5, 2026. 52.5% fewer hallucinations vs 5.3, 30% shorter responses, first Instant-class High-capability rating on cybersec/bio-chem.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":40,"comms":30,"multimodal":30,"agentic":10}},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","provider":"Google","tier":"frontier","released":"2026-02-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":54.2,"swe_verified":78,"livecodebench":76,"aime":85,"tau2":78,"api_in":2,"api_out":12,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":119,"subscription":"Gemini Advanced $20 · Ultra $100","notes":"Released Feb 19, 2026 (preview). Pricing doubles above 200K input tokens. 50% batch discount.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":39,"reasoning":41,"knowledge":49,"comms":41,"multimodal":50,"agentic":29}},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":58.7,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78,"agents_last_exam":null,"quality_vs_fable":81.6,"api_in":1.5,"api_out":7.5,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":304,"subscription":"Gemini API · AI Studio · Gemini app · Antigravity","notes":"GA stable ID gemini-3.6-flash. Google reports fewer tool calls and 17% fewer output tokens than 3.5 Flash on the AA Index workload; Computer Use is preview. Artificial Analysis measured about 304 output tok/s at high thinking.","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":42,"reasoning":38,"knowledge":45,"comms":40,"multimodal":47,"agentic":42}},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","provider":"Google","tier":"legacy","released":"2026-05-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55.1,"swe_verified":78,"livecodebench":70,"aime":75,"tau2":68,"aaii_v4_1":50,"swe_bench_pro":55.1,"deep_swe_v1_1":37,"terminal_bench_2_1":76.2,"agents_last_exam":null,"quality_vs_fable":76.1,"api_in":1.5,"api_out":9,"api_cache_hit":0.375,"batch_discount":50,"tok_per_sec":165,"subscription":"Free CLI: 1000 req/day at 1M context","notes":"Shipped GA at Google I/O May 19, 2026. Still offered, but Gemini 3.6 Flash is the recommended migration target with stronger agentic results and lower output-token pricing.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":30,"reasoning":30,"knowledge":40,"comms":30,"multimodal":40,"agentic":20}},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":54.2,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":36,"swe_bench_pro":54.2,"deep_swe_v1_1":null,"terminal_bench_2_1":54,"agents_last_exam":null,"quality_vs_fable":65.9,"api_in":0.3,"api_out":2.5,"api_cache_hit":0.03,"batch_discount":50,"tok_per_sec":490,"subscription":"Gemini API · AI Studio · Gemini app rollout","notes":"GA stable ID gemini-3.5-flash-lite. Google positions it for high-volume subagents, document parsing and structured extraction; Artificial Analysis measured about 490 output tok/s. Computer Use availability differs across Google documentation and should be validated per API surface.","reasoning_capable":true,"effort_default":"minimal","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":32,"reasoning":28,"knowledge":34,"comms":32,"multimodal":38,"agentic":35}},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash Lite","provider":"Google","tier":"legacy","released":"2026-03-10","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":22,"swe_verified":65,"livecodebench":60,"aime":68,"tau2":58,"api_in":0.25,"api_out":1.5,"api_cache_hit":0.0625,"batch_discount":50,"tok_per_sec":220,"subscription":"Free CLI + AI Studio","notes":"Lite tier; 1M context retained.","reasoning_capable":true,"effort_default":"off","effort_levels":["off","low","medium"],"cache_discount":0.25,"capability_levels":{"coding":22,"reasoning":22,"knowledge":30,"comms":25,"multimodal":35,"agentic":15}},{"id":"deepseek-v4-pro","name":"DeepSeek V4-Pro","provider":"DeepSeek","tier":"frontier","released":"2026-04-24","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":55.4,"swe_verified":80.6,"livecodebench":93.5,"aime":88,"tau2":78,"api_in":0.435,"api_out":0.87,"api_cache_hit":0.043,"batch_discount":null,"tok_per_sec":60,"subscription":"API only","notes":"MIT-licensed. Permanent pricing May 22, 2026 — Opus-class quality at ~1/10 cost. 1.6T/49B MoE, 1M context.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.0083,"capability_levels":{"coding":47,"reasoning":40,"knowledge":38,"comms":30,"multimodal":8,"agentic":28}},{"id":"deepseek-v4-flash","name":"DeepSeek V4-Flash","provider":"DeepSeek","tier":"fast","released":"2026-07-31","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"tau2":72,"terminal_bench_2_1":82.7,"agents_last_exam":25.2,"api_in":0.14,"api_out":0.28,"api_cache_hit":0.0028,"batch_discount":null,"tok_per_sec":90,"subscription":"API + open weights","notes":"V4-Flash-0731 public beta; 284B/13B active MoE, 1M context and 384K max output. DeepSeek reports Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in its unreleased minimal harness; internal DSBench results are not normalized here. A future 2x peak-hours rate is announced, but no effective date is published.","reasoning_capable":true,"effort_default":"think","effort_levels":["non-think","think"],"cache_discount":0.02,"capability_levels":{"coding":40,"reasoning":30,"knowledge":30,"comms":25,"multimodal":5,"agentic":25}},{"id":"minimax-m2.7","name":"MiniMax M2.7","provider":"MiniMax","tier":"frontier","released":"2026-03-18","license":"open-weight","jurisdiction":"China","context":205000,"context_label":"205K","swe_pro":40,"swe_verified":74,"livecodebench":72,"aime":80,"tau2":73,"api_in":0.3,"api_out":1.2,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":70,"subscription":"API + open-weight","notes":"Current flagship reasoner; MoE 230B/10B active.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.52,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":20,"agentic":20}},{"id":"grok-4.3","name":"Grok 4.3","provider":"xAI","tier":"frontier","released":"2026-05-06","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":45,"swe_verified":80,"livecodebench":80,"aime":88,"tau2":76,"api_in":1.25,"api_out":2.5,"api_cache_hit":0.31,"batch_discount":null,"tok_per_sec":75,"subscription":"SuperGrok $30 · Heavy $300","notes":"Current xAI flagship — aggressive pricing for frontier tier. Hybrid reasoning. Multimodal text+image.","reasoning_capable":true,"effort_default":"reasoning-on","effort_levels":["off","on"],"cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":44,"knowledge":42,"comms":34,"multimodal":34,"agentic":34}},{"id":"grok-4.1-fast","name":"Grok 4.1 Fast","provider":"xAI","tier":"fast","released":"2026-03-20","license":"proprietary","jurisdiction":"US","context":2000000,"context_label":"2M","swe_pro":30,"swe_verified":70,"livecodebench":65,"aime":75,"tau2":65,"api_in":0.2,"api_out":0.5,"api_cache_hit":0.05,"batch_discount":null,"tok_per_sec":140,"subscription":"SuperGrok $30","notes":"Cheapest large-context model on market. 2M context, $0.05/Mtok cached input.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.25,"capability_levels":{"coding":28,"reasoning":28,"knowledge":30,"comms":26,"multimodal":26,"agentic":22}},{"id":"mistral-medium-3.5","name":"Mistral Medium 3.5","provider":"Mistral","tier":"frontier","released":"2026-04-29","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":42,"swe_verified":77.6,"livecodebench":75,"aime":75,"tau2":65,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":75,"subscription":"Le Chat Pro €20","notes":"EU jurisdiction. 128B dense, Apache 2.0. Strongest non-Chinese open-weight coding agent. Vibe agents (GitHub/Linear/Jira/Sentry integrations). Pricing not verified; Medium 3 was $0.40/$2.00.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":40,"reasoning":30,"knowledge":40,"comms":30,"multimodal":10,"agentic":20}},{"id":"mistral-large-3","name":"Mistral Large 3","provider":"Mistral","tier":"frontier","released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"tau2":70,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":65,"subscription":"API + Le Chat","notes":"Cheapest premium output price in market among Western frontier. 675B/41B MoE, Apache 2.0.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":38,"reasoning":36,"knowledge":40,"comms":36,"multimodal":20,"agentic":28}},{"id":"subq-1m-preview","name":"SubQ 1M-Preview","provider":"SubQ","tier":"frontier","released":"2026-05-15","license":"proprietary","jurisdiction":"US","context":12000000,"context_label":"12M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":1,"api_out":5,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":80,"subscription":"API preview","notes":"First commercial subquadratic (non-transformer) LLM. ~1/5 frontier cost on long-context tasks. Capability rating estimated — public benchmarks pending.","reasoning_capable":null,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":30,"reasoning":32,"knowledge":36,"comms":30,"multimodal":5,"agentic":28}},{"id":"nemotron-3-ultra","name":"Nvidia Nemotron 3 Ultra","provider":"Nvidia","tier":"frontier","released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"tau2":76,"api_in":1.5,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":300,"subscription":"build.nvidia.com (closed API tier) + open weights","notes":"550B/55B MoE, hybrid Mamba-Transformer. ~300 tok/s. Open weights also available — see self_hosting. Capability levels are best estimates pending independent benchmarks.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":42,"reasoning":42,"knowledge":42,"comms":34,"multimodal":10,"agentic":34}},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","provider":"Swiss AI Initiative","tier":"fast","released":"2026-06-15","license":"Apache 2.0","jurisdiction":"Switzerland","context":4096,"context_label":"4K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Self-hosted weights","notes":"Largest newly released Apertus v1.1 distilled checkpoint. Dense 4.6B storage / 3.8B compute parameters, trained on 1.7T tokens, supports 1,811 languages, and ships in BF16 plus server and Apple-oriented quantizations. No comparable coding-agent, throughput, or API-price evidence is published.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null},{"id":"inkling","name":"Inkling","provider":"Thinking Machines Lab","tier":"frontier","released":"2026-07-15","license":"Apache 2.0","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Open weights (Hugging Face); fine-tuning and inference via the Tinker platform and third-party providers","notes":"975B-parameter multimodal MoE (41B active), 66-layer decoder-only transformer, 1M-token context. Text/image/audio input, text-only output. No first-party per-token API pricing published -- Thinking Machines monetizes via the Tinker fine-tuning platform rather than metered inference. A smaller Inkling-Small (12B active) companion model was released alongside it. Benchmark results are published on the model card across reasoning, agentic, coding, factuality, vision, audio, and safety categories but are not yet normalized into this roster's comparable benchmark set.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null}],"subscriptions":[{"provider":"Anthropic","tier":"Pro","price_usd":20,"limits":"~45 messages / 5h on Sonnet · limited Opus","models":"Sonnet 4.6, Haiku 4.5, limited Opus 4.8/4.7","features":"Chat only (Claude Code removed April 2026). From June 15, 2026 split into Chat pool + Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 5×","price_usd":100,"limits":"5× Pro quotas · ~225 msg/5h Sonnet · expanded Opus","models":"Full Opus 4.8/4.7 · Sonnet 4.6 · Haiku 4.5","features":"Cache reads included flat-rate. From June 15, 2026: Chat pool + separate Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 20×","price_usd":200,"limits":"20× Pro quotas","models":"All","features":"For heavy Opus users. From June 15, 2026: Chat + Agent SDK credit pools."},{"provider":"OpenAI","tier":"Plus","price_usd":20,"limits":"~80 GPT-5.4 msg/3h","models":"GPT-5.4 (limited) · GPT-5.4 Mini · o-series","features":"ChatGPT · GPTs · Codex CLI 30-150 tasks/5h"},{"provider":"OpenAI","tier":"Pro 5×","price_usd":100,"limits":"5× Plus quotas (new tier April 2026)","models":"GPT-5.4 Thinking unlimited","features":"Released as Anthropic Max competitor"},{"provider":"OpenAI","tier":"Pro","price_usd":200,"limits":"Effectively unlimited","models":"All including o3-pro","features":"Original premium tier"},{"provider":"Google","tier":"Gemini Advanced","price_usd":20,"limits":"Generous, soft caps","models":"Gemini 3.1 Pro · Gemini 3.6 Flash · Gemini 3.5 Flash-Lite","features":"Workspace integration · 1M context"},{"provider":"Google","tier":"Ultra","price_usd":100,"limits":"Higher quotas + Veo video","models":"All Gemini + research preview","features":"Veo 3 video · Project Mariner"},{"provider":"Mistral","tier":"Le Chat Pro","price_usd":22,"limits":"Generous","models":"Mistral Large 3 · Codestral","features":"EU jurisdiction"},{"provider":"xAI","tier":"SuperGrok","price_usd":30,"limits":"Generous","models":"Grok 4 · Grok 4 Heavy","features":"X integration"},{"provider":"DeepSeek","tier":"API only","price_usd":null,"limits":"Pay per token","models":"V4-Pro · V4-Flash","features":"Cheapest frontier API ($0.435/$0.87 permanent since May 22, 2026)"}],"agent_policies":[{"provider":"Anthropic","subscription_automated":"Prohibited","enforcement":"Active (OAuth blocked Apr 4)","first_party_exception":"`claude -p` pipe mode and Claude Code itself","api_required_for_automation":true,"cite":[35]},{"provider":"OpenAI","subscription_automated":"Prohibited (ToS)","enforcement":"Currently tolerated","first_party_exception":"Codex CLI uses subscription quota","api_required_for_automation":false},{"provider":"Google","subscription_automated":"Allowed in CLI","enforcement":"—","first_party_exception":"Gemini CLI free tier 1000 req/day","api_required_for_automation":false},{"provider":"Mistral","subscription_automated":"Allowed","enforcement":"—","first_party_exception":"—","api_required_for_automation":false}],"harnesses":[{"id":"claude-code","name":"Claude Code","vendor":"Anthropic","license":"proprietary","category":"CLI + IDE","stars":49000,"providers":["Anthropic"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":true,"computer_use":true,"lsp":true,"git":true,"memory":"CLAUDE.md","sandbox":"local","swe_pro":46,"pricing":"Subscription Pro/Max","sweet_spot":"SubagentStop hooks, /plugin list, requiredMinimumVersion managed setting, MCP fixes, Agent Teams.","stumbles":"Anthropic-only. Removed from standard Pro tier April 2026 — push to Max.","cite":[23,1]},{"id":"opencode","name":"OpenCode","vendor":"Anomaly","license":"MIT","category":"CLI + ACP","stars":188119,"providers":["75+: Anthropic*, OpenAI, Google, Mistral, Kimi, GLM, Ollama, LM Studio"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.18.4 (Jul 20, 2026): adaptive Kimi thinking controls, provider-defined reasoning options, restored Azure endpoints, and a rewritten desktop prompt input.","stumbles":"Anthropic OAuth blocked April 4 — must use API key.","cite":[24,35,72]},{"id":"codex-cli","name":"Codex CLI","vendor":"OpenAI","license":"Apache 2.0","category":"CLI + macOS app","stars":100232,"providers":["OpenAI"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"cloud","swe_pro":56.8,"pricing":"Plus $20 / Pro $200","sweet_spot":"v0.144.6 (Jul 18, 2026): refreshed GPT-5.6 Sol/Terra/Luna bundled instructions and corrected Codex context-window metadata to 272K.","stumbles":"OpenAI-only. No MCP, no hooks. Tightly coupled to apply_patch tool.","cite":[25,36,70]},{"id":"gemini-cli","name":"Gemini CLI","vendor":"Google","license":"Apache 2.0","category":"CLI","stars":106096,"providers":["Google"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"GEMINI.md","sandbox":"local","swe_pro":null,"pricing":"Free 1000 req/day","sweet_spot":"v0.51.0 (Jul 16, 2026): hardened sensitive-path and symlink handling, read-only macOS sandbox git config, and modern-model escape-sequence fixes.","stumbles":"Sunsetting to Antigravity CLI for free tier on June 18, 2026; paid Gemini/Enterprise keys retain access.","cite":[26,71]},{"id":"aider","name":"Aider","vendor":"paul-gauthier","license":"Apache 2.0","category":"CLI","stars":32000,"providers":["Anthropic*","OpenAI","Google","Ollama","100+"],"mcp":false,"skills":false,"hooks":false,"subagents":false,"voice":true,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"CONVENTIONS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"Mature, lightweight, voice-input native. Pair-programmer mode. Active, last commit Mar 2026. 44k stars.","stumbles":"No MCP, no hooks. Less ambitious than newer harnesses.","cite":[27]},{"id":"cline","name":"Cline","vendor":"cline-bot","license":"Apache 2.0","category":"VS Code extension","stars":64886,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".clinerules","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v4.0.10 (Jul 20, 2026): current release adds telemetry for consecutive-mistake-limit events; multi-editor and CLI surfaces remain available.","stumbles":"Anthropic OAuth blocked. Can be expensive on long sessions.","cite":[28,35,73]},{"id":"roo-code","name":"Roo Code","vendor":"RooVetGit","license":"Apache 2.0","category":"VS Code extension","stars":18000,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".roo","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v3.53.0 (Apr 23, 2026). Power-user Cline fork, model-agnostic, BYOK. Recent: GPT-5.5 via Codex, Opus 4.7 on Vertex, checkpoint nav.","stumbles":"Same OAuth situation. Configuration complexity.","cite":[29,35]},{"id":"cursor","name":"Cursor","vendor":"Anysphere","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".cursorrules","sandbox":"local","swe_pro":null,"pricing":"Hobby free · Pro $20 · Pro+ $60 · Ultra $200 · Teams $40/user","sweet_spot":"Cursor 3.5 (May 20, 2026): Cloud Agents (isolated VMs, multi-repo), Composer 2.5, Agents Window, parallel subagents.","stumbles":"Closed source. Lock-in. Subscription required for serious use.","cite":[31]},{"id":"windsurf","name":"Windsurf / Devin Desktop","vendor":"Cognition","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"windsurfrules","sandbox":"local","swe_pro":null,"pricing":"Pro $20 · Max $200","sweet_spot":"Renamed to Devin Desktop, Agent Command Center kanban, embedded Devin cloud agent, SWE-1.6 model, multi-agent + worktrees.","stumbles":"Pro $20 (was $15); smaller ecosystem than Cursor.","cite":[32]},{"id":"goose","name":"Goose","vendor":"Block","license":"Apache 2.0","category":"Desktop + CLI","stars":51387,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":false,"git":true,"memory":".goosehints","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.43.0 (Jul 14, 2026): per-message token/cost/TTFT/tok-s metrics, ACP reconnection, GPT-5.6 support, dynamic Ollama Cloud discovery, and expanded providers.","stumbles":"Anthropic OAuth blocked. Less mindshare than OpenCode.","cite":[30,35,74]},{"id":"omo","name":"OMO (Multi-model orchestrator)","vendor":"community","license":"MIT","category":"Multi-agent orchestrator","stars":54000,"providers":["multi-provider"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free","sweet_spot":"Rebrand from oh-my-opencode; multi-model orchestration. Wraps Claude Code, OpenCode, Codex, Kimi K2, DeepSeek V4, Gemini CLI.","stumbles":"Niche. Steep learning curve."},{"id":"hermes","name":"Hermes Agent","vendor":"Nous Research","license":"MIT","category":"Self-hosted multi-platform agent","stars":null,"providers":["OpenRouter-style multi-model"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":true,"remote":true,"computer_use":true,"lsp":false,"git":true,"memory":"persistent memory + auto-gen skills","sandbox":"local/docker/ssh/singularity/modal","swe_pro":null,"pricing":"Free (self-hosted)","sweet_spot":"v0.16.0. Persistent memory + auto-generated skills — learns your projects. Bridges Telegram/Discord/Slack/WhatsApp/Signal/Email/CLI. Natural-language cron for unattended runs. Parallel isolated subagents. Web search, browser automation, vision, image-gen, TTS.","stumbles":"Not coding-specialised — general autonomous agent. Setup requires self-host script. No first-party SWE benchmark.","cite":[]},{"id":"github-copilot-cli","name":"GitHub Copilot CLI","vendor":"GitHub","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["GitHub Models"],"mcp":false,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":true,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Bundled with Copilot Business/Enterprise","sweet_spot":"GA Feb 25, 2026. Specialized sub-agents (Explore, Task, Code Review, Plan), background delegation, autopilot.","stumbles":"Bundled-only — no standalone tier. GitHub-centric."},{"id":"amp-cli","name":"Amp CLI","vendor":"Sourcegraph","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["Multi"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Subscription (Sourcegraph)","sweet_spot":"Spun out as standalone company 2026. Runs as sidebar agent inside Zed via Terminal Threads.","stumbles":"Early standalone phase."},{"id":"zed","name":"Zed","vendor":"Zed Industries","license":"proprietary","category":"IDE","stars":null,"providers":["15 LLM providers + MCP"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Personal free (2k predictions) · Pro $10 · Business $30/seat","sweet_spot":"Rust-native editor with first-class agent panel + ACP host. Terminal Threads run Claude Code/Amp inline.","stumbles":"Editor first; agent layer still maturing."},{"id":"continue-dev","name":"Continue","vendor":"Continue","license":"Apache 2.0","category":"IDE extension + CLI","stars":null,"providers":["multi-provider"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".continue","sandbox":"local","swe_pro":null,"pricing":"Solo $0 · Team/Company ~$10/dev/mo","sweet_spot":"Agent mode plan+execute. Continuous AI, Mission Control, shared PR/ticket workflows.","stumbles":"Newer agent features still stabilizing across providers."}],"self_hosting":{"hardware_options":[{"id":"vast-2xa6000","name":"Vast.ai 2× A6000","vram_gb":96,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Reference cloud-GPU setup for this dashboard. No upfront capex. Docker templates, SSH/Cloudflare Zero Trust. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-3xa6000","name":"Vast.ai 3× A6000","vram_gb":144,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Headroom for Qwen 3 235B-A22B Q6_K. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-4xa6000","name":"Vast.ai 4× A6000","vram_gb":192,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Llama 3.1 405B Q4 territory. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"mbp-14-m4pro-64","name":"MacBook Pro 14\" M4 Pro 64GB","vram_gb":64,"cost_label":"~$3,200 capex","cost_per_hour":null,"type":"local","notes":"MoE sweet spot. 273 GB/s memory bandwidth."},{"id":"mbp-16-m5max-128","name":"MacBook Pro 16\" M5 Max 128GB","vram_gb":128,"cost_label":"~$5,500 capex","cost_per_hour":null,"type":"local","notes":"Best portable inference. ~545 GB/s bandwidth."},{"id":"mba-15-32","name":"MacBook Air 15\" 32GB","vram_gb":32,"cost_label":"~$1,900 capex","cost_per_hour":null,"type":"local","notes":"Hard 32GB ceiling. Limited to ~14B dense or 26B MoE Q4."},{"id":"contabo-2xl40s","name":"Contabo 2 x L40S","provider":"Contabo","vram_gb":96,"cost_label":"€1,502/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"EU-oriented fixed monthly configuration; 64 vCPU, 213 GB RAM, 3.5 TB storage, and 15 TB bandwidth are listed for the 2-GPU tier."},{"id":"contabo-1xh200","name":"Contabo 1 x H200 NVL","provider":"Contabo","vram_gb":141,"cost_label":"€2,149/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"Fixed monthly large-memory option. Confirm location and availability before treating it as a sovereignty or latency fit."},{"id":"infomaniak-1xl40s","name":"Infomaniak 1 x L40S","provider":"Infomaniak","vram_gb":48,"cost_label":"Live calculator / availability validation","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"calculator_or_request","price_checked":"2026-07-19","source":"https://www.infomaniak.com/en/hosting/public-cloud/prices","notes":"Swiss OpenStack option with dedicated GPU access and usage billing. Public pages list L40S availability but do not expose a stable crawlable SKU price; validate availability for the selected region."},{"id":"hyperstack-1xh200","name":"Hyperstack 1 x H200 SXM","provider":"Hyperstack","vram_gb":141,"cost_label":"$3.50/hr on demand","cost_per_hour":3.5,"type":"cloud","show_in_fit":false,"price_status":"published_on_demand","price_checked":"2026-07-19","source":"https://www.hyperstack.cloud/","notes":"Minute-accurate on-demand billing; reservation pricing starts at $2.45/hr. Validate region, storage, and availability."}],"models":[{"id":"gemma-4-26b-moe","name":"Gemma 4 26B-A4B MoE","params_total":26,"params_active":4,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":35,"livecodebench":77,"aime":88,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-3xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-4xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":22,"tok_per_sec":75},"mbp-16-m5max-128":{"quant":"BF16","vram_used":52,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"MoE — only 4B active per token. Faster than dense models 5× its size."},{"id":"gemma-4-31b-dense","name":"Gemma 4 31B Dense","params_total":31,"params_active":31,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":38,"livecodebench":80,"aime":89,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-3xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-4xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":19,"tok_per_sec":28},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":26,"tok_per_sec":38},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Highest quality open Gemma. Slower per-token (full 31B active)."},{"id":"qwen-3.6-plus","name":"Qwen 3.6 Plus","params_total":397,"params_active":17,"released":"2026-04-11","license":"Apache 2.0","jurisdiction":"China","swe_pro":50,"livecodebench":71.4,"aime":87,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":38},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":35},"vast-4xa6000":{"quant":"Q8_0","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1M context. Top open agentic coder. April 11 release.","swe_verified":68.2},{"id":"llama-4-maverick","name":"Llama 4 Maverick","params_total":400,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":42,"livecodebench":70,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":28},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":25},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"400B / 17B active MoE. 1M context. Strong MMLU-Pro (80.5%) but coding behind Chinese labs.","swe_verified":72},{"id":"minimax-m2.5-open","name":"MiniMax M2.5 (open weights)","params_total":456,"params_active":46,"released":"2026-01-20","license":"open-weight","jurisdiction":"China","swe_pro":40,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Eclipsed by DeepSeek V4 and Qwen 3.6 Plus. China jurisdiction. Maintained for niche workloads."},{"id":"deepseek-v4-pro-open","name":"DeepSeek V4-Pro (open weights)","params_total":1600,"params_active":49,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":55.4,"livecodebench":93.5,"aime":90,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q2_K","vram_used":188,"tok_per_sec":14},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1.6T total / 49B active MoE. MIT license. Strongest open coder. Requires very large multi-GPU deployment."},{"id":"llama-4-scout","name":"Llama 4 Scout","params_total":109,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":36,"swe_verified":68,"livecodebench":66,"aime":78,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-3xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-4xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":32,"tok_per_sec":22},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":45,"tok_per_sec":32},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"10M token context (longest in any open model). 109B / 17B active MoE."},{"id":"kimi-k2.6","name":"Kimi K2.6","params_total":235,"params_active":21,"released":"2026-05-01","license":"Modified MIT","jurisdiction":"China","swe_pro":47,"swe_verified":75,"livecodebench":78,"aime":84,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":36},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":32},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Best open-weight for sub-agent fan-out. Built for harness-driven parallel pipelines. Chinese-trained."},{"id":"glm-5.1","name":"GLM 5.1","params_total":358,"params_active":32,"released":"2026-04-22","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":76,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT license — rare among open frontier models besides DeepSeek. Strong for enterprise fine-tuning."},{"id":"mistral-small-4","name":"Mistral Small 4","params_total":24,"params_active":24,"released":"2026-04-18","license":"Apache 2.0","jurisdiction":"EU","swe_pro":28,"swe_verified":60,"livecodebench":58,"aime":68,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-3xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-4xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":48,"tok_per_sec":38},"mbp-16-m5max-128":{"quant":"BF16","vram_used":48,"tok_per_sec":55},"mba-15-32":{"quant":"Q4_K_M","vram_used":14,"tok_per_sec":30}},"notes":"6.5B effective parameters. EU jurisdiction. Best on-device option."},{"id":"nemotron-3-ultra-550b-a55b-moe","name":"Nvidia Nemotron 3 Ultra","params_total":550,"params_active":55,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q3_K_M","vram_used":180,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Hybrid Mamba-Transformer. 1M context. NVIDIA Open Model License. Computex June 4 launch."},{"id":"nemotron-3-nano-30b-a3b","name":"Nvidia Nemotron 3 Nano 30B-A3B","params_total":30,"params_active":3,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":30,"swe_verified":70,"livecodebench":68,"aime":76,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":24,"tok_per_sec":70},"mbp-16-m5max-128":{"quant":"BF16","vram_used":60,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":17,"tok_per_sec":50}},"notes":"Hybrid Mamba-Transformer Nano variant. Open weights."},{"id":"nemotron-nano-9b-v2","name":"Nvidia Nemotron Nano 9B v2","params_total":9,"params_active":9,"released":"2026-04-12","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":22,"swe_verified":60,"livecodebench":55,"aime":65,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-3xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-4xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":18,"tok_per_sec":90},"mbp-16-m5max-128":{"quant":"BF16","vram_used":18,"tok_per_sec":120},"mba-15-32":{"quant":"Q4_K_M","vram_used":6,"tok_per_sec":65}},"notes":"Dense 9B. Strong instruction following at edge sizes."},{"id":"kimi-k2.6-1t","name":"Kimi K2.6 (1T MoE)","params_total":1000,"params_active":32,"released":"2026-04-20","license":"Modified MIT","jurisdiction":"China","swe_pro":58.6,"swe_verified":80.2,"livecodebench":82,"aime":88,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Top open intelligence. 80.2% SWE-Verified, 58.6% SWE-Pro, AAII 54. 262K context. Modified MIT."},{"id":"glm-4.6","name":"GLM 4.6","params_total":355,"params_active":32,"released":"2025-09-15","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":75,"aime":80,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":28},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":24},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 200K context. Z.ai release."},{"id":"qwen3.6-35b-a3b","name":"Qwen 3.6 35B-A3B","params_total":35,"params_active":3,"released":"2026-04-16","license":"Apache 2.0","jurisdiction":"China","swe_pro":38,"swe_verified":74,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":22,"tok_per_sec":80},"mbp-16-m5max-128":{"quant":"BF16","vram_used":70,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"Apache 2.0 MoE — strong $/quality for fast bulk inference."},{"id":"cohere-command-a-plus","name":"Cohere Command A+","params_total":218,"params_active":28,"released":"2026-05-22","license":"CC-BY-NC + commercial","jurisdiction":"Canada","swe_pro":42,"swe_verified":75,"livecodebench":70,"aime":78,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":30},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":26},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"First open-weights Cohere release in &gt;1 year. Sparse-MoE multimodal."},{"id":"mistral-large-3-open","name":"Mistral Large 3 (open weights)","params_total":675,"params_active":41,"released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q3_K_M","vram_used":140,"tok_per_sec":22},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":180,"tok_per_sec":20},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"675B/41B active MoE. Apache 2.0. EU jurisdiction."},{"id":"deepseek-v4-flash-open","name":"DeepSeek V4-Flash (open weights)","params_total":284,"params_active":13,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":45},"vast-3xa6000":{"quant":"Q6_K","vram_used":115,"tok_per_sec":40},"vast-4xa6000":{"quant":"Q8_0","vram_used":150,"tok_per_sec":34},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":20},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 1M context, 384K max output. Cheap large-context open option."}],"frameworks":[{"name":"llama.cpp","best_for":"Mac (MLX), broad GGUF support","notes":"Best Apple Silicon performance via Metal."},{"name":"vLLM","best_for":"Multi-GPU servers, throughput","notes":"Production serving. Tensor parallelism."},{"name":"Ollama","best_for":"Easiest setup, dev workflow","notes":"Wrapper around llama.cpp. One-line model pull."},{"name":"MLX","best_for":"Apple Silicon native","notes":"Apple's framework. Best M-series perf."}]},"strategy":{"current_recommendation":{"label":"Risk-tiered GPT-5.6 + Fable stack","monthly_usd":null,"components":["GPT-5.6 Luna low — high-volume subagents and routine transformations","Gemini 3.5 Flash-Lite minimal — throughput-first extraction, parsing and parallel subagents","Gemini 3.6 Flash medium — Google-first coding, computer-use and multimodal agent loops","GPT-5.6 Terra medium — default engineering, review and documentation","GPT-5.6 Sol high or Claude Fable 5 high — escalation for hard, long-horizon tasks","Qwen 3.6 35B-A3B or DeepSeek V4-Flash — local route for privacy-sensitive bulk work"],"rationale":"Route by task risk instead of one subscription. Terra and Luna now retain strong coding-agent quality at substantially lower list price; Fable remains the long-horizon option only where its 30-day retention requirement is acceptable. Monthly cost is workload-dependent and must be computed from measured token volume."},"alternatives":[{"label":"Dual subscription (Claude Max + OpenAI Pro)","monthly_usd":230,"rationale":"Adds GPT-5.4 SWE-bench Pro lead and Codex CLI cloud sandbox. Worth $100/mo only if you frequently hit hard issues where Sonnet 4.6 plateaus.","verdict":"Defer until you have measured Sonnet plateau frequency for 30 days."},{"label":"API-only (no subscriptions)","monthly_usd":200,"rationale":"Pure pay-per-use. Maximum flexibility. Loses the subscription cache advantage — same workload costs 1.5–5× more for power users.","verdict":"Worse economics for your usage volume. Skip."},{"label":"Self-hosted maximalist","monthly_usd":80,"rationale":"Vast.ai 24/7 with Gemma 4 + Qwen 3 + occasional API top-up for frontier-only tasks.","verdict":"Cheapest if quality plateau at ~88% of Opus is acceptable. Operational overhead is real."}],"routing":[{"tier":"Bulk (70%)","use_for":"Classification, simple edits, triage, log parsing","preferred":"Gemini 3.5 Flash-Lite minimal · GPT-5.6 Luna low · Qwen 3.6 35B-A3B when local","cost_label":"Gemini $0.30/$2.50; Luna $0.20/$1.20 Mtok before cache, batch and reasoning tokens"},{"tier":"Mid (25%)","use_for":"Multi-file edits, code review, refactors, docs","preferred":"GPT-5.6 Terra medium · Gemini 3.6 Flash medium for Google-first or multimodal work","cost_label":"Terra $2/$12; Gemini $1.50/$7.50 Mtok before cache, batch and reasoning tokens"},{"tier":"Premium (5%)","use_for":"Architecture, hard debugging, long-context refactors","preferred":"GPT-5.6 Sol high · Claude Fable 5 high for long-horizon autonomy","cost_label":"Sol $5/$30; Fable $10/$50 Mtok before reasoning tokens"}],"open_questions":["Will the Nvidia Nemotron Coalition (Mistral, Cursor, Black Forest Labs, Thinking Machines) actually deliver a frontier-open consortium, or fragment within a quarter?","How does the June 15 Anthropic billing restructure (Chat pool + Agent SDK credit pool) change Max economics for agentic workloads?","Does Gemini 3.6 Flash’s lower output price and stronger agentic performance displace 3.1 Pro for most coding workflows before Gemini 3.5 Pro arrives?","DeepSeek V4-Pro permanent pricing ($0.435/$0.87) — does it force closed providers (OpenAI, Anthropic) to cut headline prices?","SubQ 1M-Preview's subquadratic architecture — does the next year see other commercial non-transformer LLMs?"],"task_fit":{"description":"For each task type, one recommended pick per provider drawn from the full roster. Click any cell to see runners-up that the daily briefing also considered. Burn × is derived from the matrix and recomputes when you change the reference dropdown.","rows":[{"task":"Tiny edits / boilerplate","description":"Single-file syntax fixes, regex replacements, formatting, docstrings, import sorting","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fixed-rate, fastest paid Anthropic option"},"openai":{"model_id":"gpt-5.5","effort":"minimal","rationale":"Cheapest reasoning setting on flagship"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Fastest Google route for cheap, high-volume pattern work"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Effectively $0/tok on existing Vast.ai infra"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ minimal","gpt-5.4-nano @ minimal"],"anthropic":["sonnet-4.6 @ low"],"google":["gemini-3.1-pro @ off"],"self_hosted":["llama-3.1-70b @ Q6_K"]}},{"task":"Normal coding task","description":"Single-function implementation, straightforward bug fixes, simple feature additions","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best $/quality on Anthropic side; cache reads included in Max 5×"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-specialised, ~⅓ burn of GPT-5.5 for near-identical SWE-Pro"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"58.7% SWE-Pro and stronger agentic coding than 3.5 Flash at lower output price"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Highest-quality open Gemma at full BF16 in 96GB"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium","gpt-5.5 @ low"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ high"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Multi-file implementation","description":"Feature spanning 3-10 files, requires understanding cross-file context","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Best style/intent understanding for code spanning files"},"openai":{"model_id":"gpt-5.5","effort":"medium","rationale":"Strong on planning; medium gives consistent multi-file edits"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Improved multi-step coding loops, fewer unwanted edits and 1M context"},"self_hosted":{"model_id":"qwen-3.6-plus","effort":null,"rationale":"Largest open MoE; strong reasoning across files"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ high","gpt-5.4 @ high"],"anthropic":["opus-4.7 @ high"],"google":["gemini-3.1-pro @ high"],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Hard debugging / architecture","description":"Race conditions, performance bottlenecks, design decisions with long-term implications","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Default Claude Code effort for Opus; best long-context retention"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"SWE-Pro leader; high effort balances depth and burn"},"google":{"model_id":"gemini-3.1-pro","effort":"high","rationale":"Preview Pro remains the deepest Google reasoning route; compare 3.6 Flash before paying the premium"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open-weight models trail frontier ~10-15pts here; not yet ready for hardest cases"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh","gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ high (lower burn)"],"google":[],"self_hosted":["qwen-3-235b-a22b — viable for some cases"]}},{"task":"Very hard autonomous repo task","description":"Overnight runs, agentic loops, tasks you don't supervise turn-by-turn","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"max","rationale":"Autonomy earns max effort's premium; no human in the loop to correct"},"openai":{"model_id":"gpt-5.5","effort":"xhigh","rationale":"Highest available reasoning budget on OpenAI side"},"google":{"model_id":null,"effort":null,"rationale":"Gemini 3.6 Flash improves agent loops but still lacks a max-equivalent tier for the hardest unsupervised runs"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Not recommended — frontier models earn their cost on the hardest tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ xhigh — better but $100/mo gating"],"anthropic":["opus-4.7 @ xhigh — if max budget is too steep"],"google":[],"self_hosted":[]}},{"task":"Code review / PR comments","description":"Reviewing diffs, suggesting improvements, finding issues without writing code yourself","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Critical thinking at moderate burn; doesn't need Opus depth"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-tuned reviewer at ⅓ burn of GPT-5.5"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Better instruction following and fewer unwanted code edits than 3.5 Flash"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Critique tasks don't need frontier; quality plateau acceptable"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ medium"],"self_hosted":["qwen-3-235b-a22b"]}},{"task":"Long-context refactor (50K+ tokens)","description":"Cross-cutting changes across a large codebase, where retention is the bottleneck","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"high","rationale":"1M context with retention that actually works; Opus's sweet spot"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"400K API context; degrades faster than Opus past 200K"},"google":{"model_id":"gemini-3.6-flash","effort":"high","rationale":"Google reports a 54% 1M-context MRCR score, roughly double 3.5 Flash"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open models cap around 200K useful context; not yet competitive"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh — if context fits in 256K"],"anthropic":["opus-4.7 @ xhigh"],"google":["gemini-3.1-pro @ high"],"self_hosted":[]}},{"task":"Subagent / delegated subtask","description":"Spawned by a main agent for focused work; quality threshold lower than user-facing","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fast, cheap, no effort knob to manage"},"openai":{"model_id":"gpt-5.4-mini","effort":"medium","rationale":"Best subagent — 94% of GPT-5.4 coding at 6× less"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Purpose-built for high-volume subagents at roughly 490 output tok/s"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Zero per-token cost for high-volume subagent loops"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ low","gpt-5.4-nano @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["llama-3.1-70b @ Q4_K_M"]}},{"task":"Documentation / explanation","description":"READMEs, code comments, tutorials, explaining existing code in prose","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best prose style on Anthropic side"},"openai":{"model_id":"gpt-5.4","effort":"medium","rationale":"Reasoning helps less here; medium-effort GPT-5.4 wins on $/quality"},"google":{"model_id":"gemini-3.6-flash","effort":"low","rationale":"Lower output price than 3.5 Flash with improved knowledge-work and document analysis"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Open models are competitive for prose tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ low"],"anthropic":["haiku-4.5"],"google":["gemini-3.5-flash-lite @ low"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Test scaffolding / fixtures","description":"Boilerplate test files, fixture generation, mocking, parameterised test suites","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Pattern-following work; doesn't need reasoning depth"},"openai":{"model_id":"gpt-5.4-mini","effort":"low","rationale":"Pattern-heavy; minimal reasoning, fast turnaround"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Lowest-latency current Google model for repetitive structured generation"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Bulk test scaffolding is exactly where self-hosted earns out"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Cybersecurity / vulnerability research","description":"Penetration testing, red-teaming, vulnerability discovery — tasks where cyber-specific safeguards apply","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Most capable Anthropic model; cyber tasks require Cyber Verification Program"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"Flagship reasoning; vetted security access via OpenAI"},"google":{"model_id":null,"effort":null,"rationale":"No cyber-specialized Gemini variant"},"self_hosted":{"model_id":"deepseek-v4-pro","effort":null,"rationale":"Open-weight without safety filtering; deploy in isolated environment"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ xhigh"],"google":[],"self_hosted":["qwen-3.6-plus","kimi-k2.6"]}}]}},"quota_burn_matrix":{"baseline_label":"selected reference model at medium effort = 1.00×","baseline_model_id":"fable-5","baseline_effort":"medium","methodology":"Burn ratios are displayed as multiples of the selected reference model at medium effort. Underlying values stored as absolute units anchored to gpt-5.5 medium = 1.00. Per-model ratios from API list pricing (firm, ±5%). Per-effort multipliers grounded in: Anthropic's published thinking-token budgets (low=skip, medium=~1k, high=5k, xhigh=10k, max=20k tokens); ArtificialAnalysis's measurement of Sonnet 4.6 max ≈ 3× Sonnet 4.5 on the Intelligence Index; nxcode.io / OpenAI guidance that xhigh ≈ 3-5× medium; ampcode's GPT-5.5 cost analysis. Per-effort multipliers are ±20%. NOTE: quality vs effort is not strictly monotonic per task — aggregate quality trends upward, but individual tasks can see high beat xhigh or medium beat high. OpenAI explicitly warns that 'high is not automatically better than medium'. Models that run in a separate preview quota bucket with no standard-pool burn are encoded as 0.00x and annotated until final token or credit rates are published.","stacking_multipliers":[{"name":"Fast mode (/fast on)","multiplier":2,"scope":"any cell"},{"name":"Cached input (long thread)","multiplier":0.6,"scope":"any cell"},{"name":"Plan mode (/plan)","multiplier":"uses high effort","scope":"overrides default"}],"openai_matrix":[{"model_id":"gpt-5.6-sol","minimal":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4,"note":"List-price blend equals the historical GPT-5.5 medium unit. Effort factors are planning estimates until measured token use is available."},{"model_id":"gpt-5.6-terra","minimal":0.18,"low":0.26,"medium":0.4,"high":0.68,"xhigh":1.08,"max":1.6,"note":"Two-fifths of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.6-luna","minimal":0.018,"low":0.026,"medium":0.04,"high":0.068,"xhigh":0.108,"max":0.16,"note":"Four percent of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.5","minimal":0.2,"low":0.5,"medium":1,"high":2,"xhigh":3.5},{"model_id":"gpt-5.5-pro","minimal":null,"low":null,"medium":6,"high":12,"xhigh":21},{"model_id":"gpt-5.4","minimal":0.1,"low":0.25,"medium":0.5,"high":1,"xhigh":1.75},{"model_id":"gpt-5.4-mini","minimal":0.011,"low":0.027,"medium":0.053,"high":0.107,"xhigh":0.187},{"model_id":"gpt-5.4-nano","minimal":0.003,"low":0.007,"medium":0.013,"high":null,"xhigh":null},{"model_id":"gpt-5.3-codex","minimal":0.07,"low":0.17,"medium":0.33,"high":0.67,"xhigh":1.17},{"model_id":"gpt-5.2","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null},{"model_id":"gpt-5.2-codex","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null}],"anthropic_matrix":[{"model_id":"fable-5","low":1.1,"medium":1.69,"high":2.87,"xhigh":4.56,"max":6.76,"note":"Always-on adaptive thinking. Values start from the $10/$50 list-price blend; effort multipliers are planning estimates rather than published token budgets."},{"model_id":"opus-4.8","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Current flagship. Same headline $5/$25. Fast Mode 3× cheaper ($10/$50) vs Opus 4.7."},{"model_id":"opus-4.7","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Legacy as of May 28. Thinking tokens: low=skip, medium=~1k, high=5k, xhigh=10k, max=20k. New tokenizer +35% tokens vs 4.6."},{"model_id":"sonnet-4.6","low":0.25,"medium":0.5,"high":1.25,"xhigh":null,"max":3.5,"note":"API default high. No xhigh tier. AA measured max ≈ 3× Sonnet 4.5 cost."},{"model_id":"haiku-4.5","low":null,"medium":0.17,"high":null,"xhigh":null,"max":null,"note":"No effort control — fixed-rate model."}],"google_matrix":[{"model_id":"gemini-3.1-pro","off":0.13,"low":0.2,"medium":0.3,"high":0.5},{"model_id":"gemini-3.5-flash","off":0.02,"low":0.03,"medium":0.05,"high":0.09},{"model_id":"gemini-3.6-flash","off":0.017,"low":0.025,"medium":0.042,"high":0.076},{"model_id":"gemini-3.5-flash-lite","off":0.005,"low":0.008,"medium":0.014,"high":0.024}],"non_reasoning_note":"Models without reasoning controls: Haiku 4.5 (Anthropic matrix above as fixed rate), GPT-5.5 Instant (ChatGPT default since May 5), Mistral Medium 3.5, MiniMax M2.7, Kimi K2.6, GLM 4.6, Qwen 3.6 variants, all Llama 4 variants, Grok 4.1 Fast, SubQ 1M-Preview — burn at fixed rate per model regardless of effort knob. DeepSeek V4-Flash-0731 now exposes thinking and non-thinking modes.","effort_quality_factors":{"minimal":0.65,"low":0.85,"medium":0.94,"high":0.98,"xhigh":1,"max":1.02,"off":0.85,"on":0.95},"quality_methodology":"Quality % = (model_swe_pro / reference_swe_pro × effort_quality_factor) × 100. Effort quality factors: minimal=0.65, low=0.85, medium=0.94, high=0.98, xhigh=1.00, max=1.02. Anchors: Anthropic's Hex measurement (low Opus 4.7 ≈ medium Opus 4.6 quality) anchors low ≈ 0.85; apiyi.com's report that max gains ~3pts over xhigh on hardest tasks anchors max=1.02; ampcode's GPT-5.5 analysis (medium captures 'most' of capability) anchors medium=0.94. IMPORTANT: these are AGGREGATE estimates. stet.sh found per-task reversals — high can beat xhigh on some tasks, medium can beat high. Treat as ±10% indicators, not precise measures.","unit_anchor":"gpt-5.5 medium","unit_anchor_note":"All burn values are stored as absolute units anchored to gpt-5.5 medium = 1.00. At render time, each cell is divided by the reference model's medium-effort cell (or its single datapoint for non-reasoning models) to produce the displayed ×-of-reference number. When the reference dropdown changes, all burn ratios recalculate against the new reference.","workload_presets":{"cold":{"label":"Cold (0% cache)","description":"One-off prompts, no context reuse. Worst-case burn.","cache_hit_rate":0},"mixed":{"label":"Mixed (40% cache)","description":"Interactive coding with some context reuse. Typical default.","cache_hit_rate":0.4},"warm":{"label":"Warm (70% cache)","description":"Sustained agentic session. AA's industry-standard 7:2:1 blend assumes this rate. danielvaughan.com's Codex CLI model uses 70%.","cache_hit_rate":0.7},"hot":{"label":"Hot (90% cache)","description":"Long warmer-pattern sessions with stable system prompts and reused context. vsits.co documented 99% achievable on Claude Code Max subscription with optimization.","cache_hit_rate":0.9}},"default_workload":"mixed","cache_methodology":"Cache discount = (cache_hit_price / input_price). Lower = better discount. Effective burn = (1 - cache_hit_rate) × raw_burn + cache_hit_rate × raw_burn × cache_discount. Per-provider cache discounts: Anthropic 10% (90% off), OpenAI 25% on flagship/10% on 5.4 family, DeepSeek V4-Pro 0.83% (99.2% off — most aggressive in industry), DeepSeek V4-Flash 2%, MiniMax 52%, Mistral &amp; Grok no published cache pricing (modeled as 100% = no discount), Google ~25% plus per-hour storage fee (not modeled). IMPORTANT CAVEAT on Anthropic subscriptions: docs say cache reads count at 10% rate, but GitHub issue anthropics/claude-code#24147 reports cache reads burning quota at full rate on Max subscriptions. Treat the displayed cache benefit as accurate for API usage; subscription quota behavior is contested. Cache writes (1.25× / 2× of input rate on Anthropic) are not modeled — amortize across many reads in steady-state."},"sources":[{"n":1,"category":"Anthropic","title":"Anthropic news","url":"https://www.anthropic.com/news"},{"n":2,"category":"Anthropic","title":"Anthropic Claude model overview","url":"https://docs.claude.com/en/docs/about-claude/models/overview"},{"n":3,"category":"Anthropic","title":"Anthropic pricing","url":"https://www.anthropic.com/pricing"},{"n":4,"category":"OpenAI","title":"OpenAI news","url":"https://openai.com/news/"},{"n":5,"category":"OpenAI","title":"OpenAI model catalog","url":"https://platform.openai.com/docs/models"},{"n":6,"category":"OpenAI","title":"OpenAI API pricing","url":"https://openai.com/api/pricing/"},{"n":7,"category":"Google","title":"Google DeepMind blog","url":"https://blog.google/technology/google-deepmind/"},{"n":8,"category":"Google","title":"Gemini API models","url":"https://ai.google.dev/gemini-api/docs/models"},{"n":9,"category":"Google","title":"Gemini API pricing","url":"https://ai.google.dev/pricing"},{"n":10,"category":"Mistral","title":"Mistral news","url":"https://mistral.ai/news/"},{"n":11,"category":"Mistral","title":"Mistral La Plateforme pricing","url":"https://mistral.ai/products/la-plateforme#pricing"},{"n":12,"category":"xAI","title":"xAI blog (Grok)","url":"https://x.ai/blog"},{"n":13,"category":"DeepSeek","title":"DeepSeek API docs","url":"https://api-docs.deepseek.com/"},{"n":14,"category":"MiniMax","title":"MiniMax news","url":"https://www.minimaxi.com/en/news"},{"n":15,"category":"Benchmarks","title":"SWE-bench (Verified)","url":"https://www.swebench.com/","note":"SWE-bench Verified is widely considered contaminated; treat with skepticism."},{"n":16,"category":"Benchmarks","title":"SWE-bench Pro leaderboard","url":"https://scale.com/leaderboard/swe_bench_pro","note":"Preferred trustworthy code-agent benchmark."},{"n":17,"category":"Benchmarks","title":"LiveCodeBench","url":"https://livecodebench.github.io/"},{"n":18,"category":"Benchmarks","title":"Artificial Analysis (cross-provider pricing + benchmarks)","url":"https://artificialanalysis.ai/"},{"n":19,"category":"Benchmarks","title":"LMArena (chatbot arena)","url":"https://lmarena.ai/"},{"n":20,"category":"Open-weight","title":"Hugging Face — trending models","url":"https://huggingface.co/models?sort=trending"},{"n":21,"category":"Open-weight","title":"Hugging Face blog","url":"https://huggingface.co/blog"},{"n":22,"category":"Open-weight","title":"r/LocalLLaMA (community signal)","url":"https://www.reddit.com/r/LocalLLaMA/"},{"n":23,"category":"Harnesses","title":"Claude Code (anthropics/claude-code)","url":"https://github.com/anthropics/claude-code"},{"n":24,"category":"Harnesses","title":"OpenCode (sst/opencode)","url":"https://github.com/sst/opencode"},{"n":25,"category":"Harnesses","title":"Codex CLI (openai/codex)","url":"https://github.com/openai/codex"},{"n":26,"category":"Harnesses","title":"Gemini CLI (google-gemini/gemini-cli)","url":"https://github.com/google-gemini/gemini-cli"},{"n":27,"category":"Harnesses","title":"Aider (Aider-AI/aider)","url":"https://github.com/Aider-AI/aider"},{"n":28,"category":"Harnesses","title":"Cline (cline/cline)","url":"https://github.com/cline/cline"},{"n":29,"category":"Harnesses","title":"Roo Code (RooVetGit/Roo-Code)","url":"https://github.com/RooVetGit/Roo-Code"},{"n":30,"category":"Harnesses","title":"Goose (block/goose)","url":"https://github.com/block/goose"},{"n":31,"category":"Harnesses","title":"Cursor changelog","url":"https://cursor.com/changelog"},{"n":32,"category":"Harnesses","title":"Windsurf changelog","url":"https://windsurf.com/changelog"},{"n":33,"category":"Hardware","title":"Vast.ai (spot GPU pricing — A6000, H100)","url":"https://vast.ai/"},{"n":34,"category":"Hardware","title":"Apple MacBook Pro (M-series)","url":"https://www.apple.com/shop/buy-mac/macbook-pro"},{"n":35,"category":"Policy","title":"Anthropic OAuth third-party restrictions (Apr 4, 2026)","url":"https://www.anthropic.com/news"},{"n":36,"category":"Policy","title":"OpenAI Codex token-based pricing migration (Apr 2, 2026)","url":"https://openai.com/news/"},{"n":37,"category":"Aggregators","title":"Hacker News (AI tags)","url":"https://news.ycombinator.com/"},{"n":38,"category":"Benchmarks","title":"AIME (math competition benchmark)","url":"https://aimeproblems.com/","note":"Used to anchor Reasoning axis ratings."},{"n":39,"category":"Benchmarks","title":"MMLU / MMLU-Pro (knowledge benchmark)","url":"https://github.com/hendrycks/test","note":"Used to anchor Knowledge axis ratings."},{"n":40,"category":"Benchmarks","title":"tau-bench (tool-use + agentic behaviour)","url":"https://github.com/sierra-research/tau-bench","note":"Used to anchor Agentic axis ratings."},{"n":41,"category":"Nvidia","title":"build.nvidia.com","url":"https://build.nvidia.com/"},{"n":42,"category":"Nvidia","title":"Hugging Face — Nvidia","url":"https://huggingface.co/nvidia"},{"n":43,"category":"Nvidia","title":"Nvidia blogs","url":"https://blogs.nvidia.com/"},{"n":44,"category":"Benchmarks","title":"SWE-bench Pro public leaderboard (Scale)","url":"https://scale.com/leaderboard/swe_bench_pro_public"},{"n":45,"category":"Benchmarks","title":"Scale Labs","url":"https://labs.scale.com/"},{"n":46,"category":"Benchmarks","title":"SWE-Rebench","url":"https://swe-rebench.com/"},{"n":47,"category":"Benchmarks","title":"Terminal-Bench 2.0","url":"https://tbench.ai/leaderboard"},{"n":48,"category":"Aggregators","title":"OpenRouter","url":"https://openrouter.ai/"},{"n":49,"category":"Aggregators","title":"LLM-Stats","url":"https://llm-stats.com/"},{"n":50,"category":"Benchmarks","title":"Artificial Analysis Intelligence Index","url":"https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index"},{"n":51,"category":"Harnesses","title":"Claude Code changelog","url":"https://code.claude.com/docs/en/changelog"},{"n":52,"category":"Harnesses","title":"Cursor pricing","url":"https://cursor.com/pricing"},{"n":53,"category":"Harnesses","title":"Windsurf changelog (Devin Desktop)","url":"https://windsurf.com/changelog"},{"n":54,"category":"Harnesses","title":"Gemini CLI changelogs","url":"https://geminicli.com/docs/changelogs"},{"n":55,"category":"Harnesses","title":"Codex CLI releases","url":"https://github.com/openai/codex/releases"},{"n":56,"category":"Google","title":"DeepMind models","url":"https://deepmind.google/models"},{"n":57,"category":"Mistral","title":"Mistral pricing","url":"https://mistral.ai/pricing"},{"n":58,"category":"Google","title":"Gemini API pricing (canonical)","url":"https://ai.google.dev/gemini-api/docs/pricing"},{"n":59,"category":"MiniMax","title":"Artificial Analysis — MiniMax M2.7","url":"https://artificialanalysis.ai/models/minimax-m2-7"},{"n":60,"category":"Moonshot","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3"},{"n":61,"category":"Moonshot","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models"},{"n":62,"category":"OpenAI","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/"},{"n":63,"category":"Hosting","title":"Runpod GPU cloud pricing","url":"https://www.runpod.io/pricing"},{"n":64,"category":"Hosting","title":"Runpod July 2024 GPU price changes","url":"https://www.runpod.io/blog/runpod-slashes-gpu-prices-more-power-less-cost-for-ai-builders"},{"n":65,"category":"Hosting","title":"Contabo GPU Cloud configurations and pricing","url":"https://contabo.com/en/gpu-cloud/"},{"n":66,"category":"Hosting","title":"Infomaniak Public Cloud pricing","url":"https://www.infomaniak.com/en/hosting/public-cloud/prices"},{"n":67,"category":"Hosting","title":"Infomaniak GPU flavor documentation","url":"https://docs.infomaniak.cloud/compute/instances/flavors/"},{"n":68,"category":"Hosting","title":"Hyperstack GPU cloud pricing","url":"https://www.hyperstack.cloud/"},{"n":69,"category":"Hosting","title":"Vast.ai marketplace pricing methodology","url":"https://docs.vast.ai/guides/instances/pricing"},{"n":70,"category":"Harnesses","title":"Codex CLI 0.144.6 release","url":"https://github.com/openai/codex/releases/tag/rust-v0.144.6"},{"n":71,"category":"Harnesses","title":"Gemini CLI 0.51.0 release","url":"https://github.com/google-gemini/gemini-cli/releases/tag/v0.51.0"},{"n":72,"category":"Harnesses","title":"OpenCode 1.18.4 release","url":"https://github.com/anomalyco/opencode/releases/tag/v1.18.4"},{"n":73,"category":"Harnesses","title":"Cline 4.0.10 release","url":"https://github.com/cline/cline/releases/tag/v4.0.10"},{"n":74,"category":"Harnesses","title":"Goose 1.43.0 release","url":"https://github.com/aaif-goose/goose/releases/tag/v1.43.0"},{"n":75,"category":"Moonshot","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing"},{"n":76,"category":"Swiss AI Initiative","title":"Apertus v1.1 4B Instruct model card","url":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"n":77,"category":"Google","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"n":78,"category":"Google","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"n":79,"category":"Google","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"n":80,"category":"Google","title":"Gemini API release notes","url":"https://ai.google.dev/gemini-api/docs/changelog"},{"n":81,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash"},{"n":82,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"},{"n":83,"category":"OpenAI","title":"GPT-5.6 launch and benchmark table","url":"https://openai.com/index/gpt-5-6/"},{"n":84,"category":"Google","title":"Gemini API Managed Agents update","url":"https://blog.google/innovation-and-ai/technology/developers-tools/expanding-managed-agents-gemini-api/"},{"n":85,"category":"Benchmarks","title":"Scale Labs SWE-bench Pro public leaderboard","url":"https://labs.scale.com/api/pdf/leaderboard/swe_bench_pro_public","note":"Public leaderboard values use their own model versions and evaluation setup; do not merge them mechanically with vendor launch tables."},{"n":86,"category":"Anthropic","title":"What's new in Claude Opus 5","url":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5","note":"Official launch specifications, pricing, availability, and migration behaviour."},{"n":87,"category":"Benchmarks","title":"Artificial Analysis — Claude Opus 5","url":"https://artificialanalysis.ai/models/claude-opus-5","note":"Independent effort-specific intelligence, latency, throughput, and price analysis."},{"n":88,"category":"Thinking Machines Lab","title":"Inkling model card","url":"https://thinkingmachines.ai/model-card/inkling/"},{"n":89,"category":"OpenAI","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra"},{"n":90,"category":"OpenAI","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna"},{"n":91,"category":"DeepSeek","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","note":"Vendor-reported benchmark results use DeepSeek Harness minimal mode at max effort; DSBench results are internal."},{"n":92,"category":"DeepSeek","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/"},{"n":93,"category":"OpenAI","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","note":"Official July 30 pricing, paid-subscription credit, and Sol API Fast mode details."},{"n":94,"category":"OpenAI","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","note":"Official August 6 ChatGPT product update; it does not change the API model contract."}],"provider_sources":{"Anthropic":[1,2,3,86,87],"OpenAI":[4,5,6,62,83,89,90,93,94],"Google":[7,8,9,58,77,78,79,80,81,82,84],"Mistral":[10,11],"xAI":[12],"DeepSeek":[13,91,92],"MiniMax":[14],"Meta":[20,21],"Alibaba":[20,21],"Moonshot":[20,21,60,61,75],"Zhipu":[20,21],"Nvidia":[41,42,43],"Cohere":[20,21],"Swiss AI Initiative":[76],"SubQ":[],"Thinking Machines Lab":[88]},"section_sources":{"models":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,60,61,62,76,77,78,79,80,81,82,86,87,89,90,91,92],"harnesses":[23,24,25,26,27,28,29,30,31,32,35,51,52,53,54,55,70,71,72,73,74],"self_hosting":[20,21,22,33,34,63,64,65,66,67,68,69,76],"strategy":[1,3,4,6,16,18,33,35,36]},"capabilities":{"schema_version":"1.0","axes":[{"key":"coding","label":"Coding","short":"Code","effort_sensitivity":0.5},{"key":"reasoning","label":"Reasoning &amp; Architecture","short":"R&amp;A","effort_sensitivity":0.6},{"key":"knowledge","label":"Knowledge &amp; Research","short":"K&amp;R","effort_sensitivity":0.05},{"key":"comms","label":"Communication &amp; Docs","short":"Comms","effort_sensitivity":0.2},{"key":"multimodal","label":"Multimodal","short":"MM","effort_sensitivity":0.05},{"key":"agentic","label":"Agentic","short":"Agent","effort_sensitivity":0.4}],"level_labels":{"coding":{"1":"Snippet","2":"Standard","3":"Cross-file","4":"Hard","5":"Frontier"},"reasoning":{"1":"Apply known","2":"Multi-step","3":"Cross-cutting","4":"Novel system","5":"Research-grade"},"knowledge":{"1":"Recall","2":"Contextual","3":"Cross-domain","4":"Frontier","5":"Original"},"comms":{"1":"Grammatical","2":"Structured","3":"Tutorial","4":"Editorial","5":"Publishable"},"multimodal":{"1":"Text only","2":"Image-in","3":"Image reasoning","4":"Video/audio","5":"Cross-modal gen"},"agentic":{"1":"Single-turn","2":"Tool calls","3":"Plan coherence","4":"Self-correcting","5":"Long-horizon"}},"focus_presets":{"balanced":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"coding-focused":{"coding":2,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"architecture-focused":{"coding":1,"reasoning":2,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"research-focused":{"coding":1,"reasoning":1,"knowledge":2,"comms":1,"multimodal":1,"agentic":1},"writing-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":2,"multimodal":1,"agentic":1},"multimodal-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":2,"agentic":1},"agentic-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":2}},"default_focus":"balanced","methodology":"1-5 levels per axis; Coding rated against SWE-Pro / SWE-Verified evidence; Reasoning against AIME / LiveCodeBench / public hard-reasoning evals; Multimodal against published modality support (image/audio/video in/out); Agentic rated holding harness constant at Claude-Code-class baseline (plan coherence, tool-call quality, self-correction, calibrated stopping, refusal hygiene); Knowledge from model-card claims + community-validated state-of-art awareness; Comms from prose-evaluation rounds + structural quality. Ratings carry citation; see src/dashboard-context.md for the full rubric.","effort_formula":"effective(axis, effort) = max(0, ceiling[axis] × (1 − effort_sensitivity[axis] × (1 − radar_effort_factor[effort]))). radar_effort_factor[medium] = 1.00 → primary polygon equals reference polygon shape at medium effort for the same model. Lower efforts shrink the polygon by sensitivity-weighted amounts; higher efforts grow it and may push vertices OUTSIDE the outer level-50 ring — the ring is a rubric marker, not a hard cap. Vertices that exceed 50 are drawn in accent-hot to signal 'boosted above medium-effort baseline'.","radar_effort_factors":{"minimal":0.5,"low":0.75,"medium":1,"high":1.08,"xhigh":1.15,"max":1.22},"scale":{"rubric_min":1,"rubric_max":50,"visual_max":60,"band_size":10,"boost_band":"51–60","interpretation":"capability_levels[axis] = the model's effective capability at MEDIUM effort. Briefings rate from 1 to 50. The radar's visual scale extends to 60 to accommodate the 'effort-boost band' (51–60) — values that fall there are formula-derived from extra reasoning effort, never directly rated. The outer rubric ring at 50 is visually emphasized as 'current frontier'.","bands":{"1-10":"Snippet","11-20":"Standard","21-30":"Cross-file","31-40":"Hard","41-50":"Frontier","51-60":"Effort-boost band (derived, not rated)"}}},"report_metrics":{"schema_version":"1.0","researched_at":"2026-08-01","currency":"USD","reference_options":["fable-5","gpt-5.6-sol","gpt-5.6-terra","gpt-5.6-luna","opus-4.8","gpt-5.5","gpt-5.5-pro","gpt-5.5-instant","kimi-k3","gemini-3.6-flash","gemini-3.5-flash-lite"],"default_reference":"fable-5","methodology":{"benchmark_scale":"Each benchmark is divided by benchmark.max_value and expressed on a 0-100 scale. The quality composite is the weighted arithmetic mean of available normalized benchmarks; it is shown only when coverage is at least 0.50.","quality_weights":{"aaii_v4_1":0.2,"coding_agent_index_v1_1":0.2,"swe_bench_pro":0.2,"deep_swe_v1_1":0.15,"terminal_bench_2_1":0.15,"agents_last_exam":0.1},"speed_score":"100 * sqrt(task_speed_index / max_task_speed_index). task_speed_index is 100 * reference_time / model_time. It is not output tokens per second and is comparable only within the cited evaluation setup.","cost_score":"10 + 90 * ln(max_blended_price / blended_price) / ln(max_blended_price / min_blended_price). Blended price is 0.30 * input_price + 0.70 * output_price per million tokens. Higher cost score is better.","scq_compound":"Geometric mean of quality_score, speed_score and cost_score. Geometric mean prevents one category from fully compensating for a weak category. Null when quality coverage is below 0.50.","capability_compound":"Weighted arithmetic mean of the six capability axes. Balanced uses weight 1 for each axis; a selected focus uses weight 2.5 for that axis and 1 for every other axis.","reference_quality":"Each model stores or derives a Fable-anchored quality value. Displayed quality = 100 * model_anchor / selected_reference_anchor, so the selected reference is always exactly 100. New-suite composites use only explicitly overlapping evaluations; legacy rows fall back to SWE-Bench Pro. Unknown evidence remains null, never zero.","burn":"Raw blended-price units are multiplied by an effort factor and cache factor, then divided by the selected reference model at medium effort. Cache factor = (1-hit_rate) + hit_rate*cache_read_ratio.","missing_data":"Keep unknown values null. Display insufficient comparable evidence rather than zero. Detailed benchmark cells remain version-specific and may be empty even when a reference-relative quality anchor exists from a separate documented comparison set."},"market_signals":[{"id":"frontier_quality","label":"Frontier quality index","explanation":"Best broad intelligence score published on Artificial Analysis Intelligence Index v4.1. Higher is better; the benchmark version must remain fixed across the series.","unit":"index","direction":"up","current":59.9,"history":[{"date":"2026-01-01","value":48.2},{"date":"2026-02-01","value":50.1},{"date":"2026-03-01","value":52.7},{"date":"2026-04-01","value":55.7},{"date":"2026-05-01","value":55.7},{"date":"2026-06-01","value":59.9},{"date":"2026-07-01","value":59.9}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"coding_quality","label":"Coding-agent frontier","explanation":"Best score on Artificial Analysis Coding Agent Index v1.1. Higher means better end-to-end coding-agent performance, not just code completion.","unit":"index","direction":"up","current":80,"history":[{"date":"2026-01-01","value":66.1},{"date":"2026-02-01","value":69.8},{"date":"2026-03-01","value":72.5},{"date":"2026-04-01","value":76.4},{"date":"2026-05-01","value":77.2},{"date":"2026-06-01","value":77.2},{"date":"2026-07-01","value":80}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"task_speed","label":"Coding task speed increase","explanation":"Fastest current end-to-end coding task-rate index, where Fable 5 is fixed at 100. A value of 320 means an estimated 3.2 times as many comparable tasks per unit time; it is not tokens per second.","unit":"Fable=100","direction":"up","current":320,"history":[{"date":"2026-01-01","value":100},{"date":"2026-02-01","value":112},{"date":"2026-03-01","value":126},{"date":"2026-04-01","value":145},{"date":"2026-05-01","value":180},{"date":"2026-06-01","value":180},{"date":"2026-07-01","value":320}],"source":"https://openai.com/index/gpt-5-6/","history_note":"Pre-July points are frozen planning estimates from prior report snapshots and should not be interpreted as one controlled longitudinal benchmark."},{"id":"blended_frontier_price","label":"Lowest frontier blended API price","explanation":"Lowest 30% input / 70% output list-price blend among models meeting the report's frontier quality floor. Lower is better; batch and caching are excluded.","unit":"$/Mtok","direction":"down","current":0.9,"history":[{"date":"2026-01-01","value":18},{"date":"2026-02-01","value":16.5},{"date":"2026-03-01","value":14},{"date":"2026-04-01","value":11.4},{"date":"2026-05-01","value":9},{"date":"2026-06-01","value":9},{"date":"2026-07-01","value":4.5},{"date":"2026-08-01","value":0.9}],"source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"open_weight_quality","label":"Open-weight quality vs selected reference","explanation":"Best open-weight model relative to the selected reference. Stored history is Fable-anchored and rebased in the browser whenever the reference changes. July uses Kimi K3's geometric mean across 14 overlapping launch-suite values; vendor-reported values require independent reproduction.","unit":"%","direction":"up","current":99.2,"history":[{"date":"2026-01-01","value":82},{"date":"2026-02-01","value":85},{"date":"2026-03-01","value":88},{"date":"2026-04-01","value":91},{"date":"2026-05-01","value":94},{"date":"2026-06-01","value":96},{"date":"2026-07-01","value":99.2}],"history_note":"Earlier points are frozen report estimates; July uses Moonshot’s K3 launch suite and is not a single controlled longitudinal benchmark.","source":"https://www.kimi.com/blog/kimi-k3"},{"id":"output_throughput","label":"Documented API output throughput","explanation":"Fastest provider-documented general API output rate in the tracked set. Unit is generated tokens per second; it is distinct from time to first token and end-to-end task speed.","unit":"tok/s","direction":"up","current":490,"history":[{"date":"2026-01-01","value":110},{"date":"2026-02-01","value":130},{"date":"2026-03-01","value":140},{"date":"2026-04-01","value":180},{"date":"2026-05-01","value":200},{"date":"2026-06-01","value":220},{"date":"2026-07-01","value":490}],"history_note":"Provider and independent figures use different serving and context conditions. The latest point is Artificial Analysis's Gemini 3.5 Flash-Lite measurement on Google's first-party API.","source":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"}],"benchmarks":[{"id":"aaii_v4_1","label":"AA Intelligence Index v4.1","max_value":59.9,"unit":"index","good_direction":"high"},{"id":"coding_agent_index_v1_1","label":"AA Coding Agent Index v1.1","max_value":80,"unit":"index","good_direction":"high"},{"id":"swe_bench_pro","label":"SWE-Bench Pro","max_value":80,"unit":"%","good_direction":"high"},{"id":"deep_swe_v1_1","label":"DeepSWE v1.1","max_value":72.7,"unit":"%","good_direction":"high"},{"id":"terminal_bench_2_1","label":"Terminal-Bench 2.1","max_value":91.9,"unit":"%","good_direction":"high","note":"Maximum is GPT-5.6 Sol Ultra; ordinary Sol scores 88.8."},{"id":"agents_last_exam","label":"Agents' Last Exam","max_value":52.7,"unit":"%","good_direction":"high"}],"model_metrics":[{"model_id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","benchmarks":{"aaii_v4_1":59.9,"coding_agent_index_v1_1":77.2,"swe_bench_pro":80,"deep_swe_v1_1":69.7,"terminal_bench_2_1":83.1,"agents_last_exam":40.5},"quality_score":94.9,"quality_coverage":1,"task_speed_index":100,"speed_score":50,"api_input":10,"api_cached_input":1,"api_output":50,"blended_price":38,"cost_score":10,"scq_compound":36.2,"speed_evidence":"Reference index; Anthropic labels comparative latency slower.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","benchmarks":{"aaii_v4_1":58.9,"coding_agent_index_v1_1":80,"swe_bench_pro":64.6,"deep_swe_v1_1":72.7,"terminal_bench_2_1":88.8,"agents_last_exam":52.7},"quality_score":95.4,"quality_coverage":1,"task_speed_index":256,"speed_score":80,"api_input":5,"api_cached_input":0.5,"api_output":30,"blended_price":22.5,"cost_score":22.6,"scq_compound":55.7,"speed_evidence":"OpenAI introduced API Fast mode on July 30, claiming up to 2.5× Standard speed at 2× price with no intelligence change; the base task-speed index remains a separate vendor comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","benchmarks":{"aaii_v4_1":55,"coding_agent_index_v1_1":77.4,"swe_bench_pro":63.4,"deep_swe_v1_1":69.6,"terminal_bench_2_1":87.4,"agents_last_exam":50.4},"quality_score":91.8,"quality_coverage":1,"task_speed_index":300,"speed_score":86.6,"api_input":2,"api_cached_input":0.2,"api_output":12,"blended_price":9,"cost_score":44.6,"scq_compound":70.8,"speed_evidence":"OpenAI reports roughly one-third of Fable 5 task time for the family coding comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","benchmarks":{"aaii_v4_1":51.2,"coding_agent_index_v1_1":74.6,"swe_bench_pro":62.7,"deep_swe_v1_1":67.2,"terminal_bench_2_1":84.7,"agents_last_exam":50.3},"quality_score":88.7,"quality_coverage":1,"task_speed_index":320,"speed_score":89.4,"api_input":0.2,"api_cached_input":0.02,"api_output":1.2,"blended_price":0.9,"cost_score":100,"scq_compound":92.6,"speed_evidence":"OpenAI calls Luna the fastest tier; 320 is a conservative planning index above the cited 300 family comparison and must not be treated as measured tok/s.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","benchmarks":{"deep_swe_v1_1":67.5,"terminal_bench_2_1":88.3},"quality_score":null,"quality_coverage":0.33,"quality_vs_fable":99.2,"task_speed_index":null,"speed_score":null,"api_input":3,"api_cached_input":0.3,"api_output":15,"blended_price":11.4,"cost_score":38.9,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time or output-throughput figure published for K3 at launch.","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"model_id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","benchmarks":{"aaii_v4_1":55.7,"coding_agent_index_v1_1":72.5,"swe_bench_pro":69.2,"deep_swe_v1_1":59,"terminal_bench_2_1":78.9,"agents_last_exam":45.2},"quality_score":87.7,"quality_coverage":1,"task_speed_index":180,"speed_score":67.1,"api_input":5,"api_cached_input":0.5,"api_output":25,"blended_price":19,"cost_score":26.7,"scq_compound":54,"speed_evidence":"Planning index based on moderate vendor latency and optional fast mode; validate in the local harness.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gemini-3.1-pro","name":"Gemini 3.1 Pro Preview","provider":"Google","benchmarks":{"aaii_v4_1":46.5,"coding_agent_index_v1_1":42.7,"swe_bench_pro":54.2,"deep_swe_v1_1":11.8,"terminal_bench_2_1":70.7,"agents_last_exam":32.1},"quality_score":59.8,"quality_coverage":1,"task_speed_index":null,"speed_score":null,"api_input":null,"api_cached_input":null,"api_output":null,"blended_price":null,"cost_score":null,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time measurement in the selected source set.","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro"},{"model_id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","benchmarks":{"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78},"quality_score":77.4,"quality_coverage":0.7,"quality_vs_fable":81.6,"task_speed_index":null,"speed_score":null,"api_input":1.5,"api_cached_input":0.15,"api_output":7.5,"blended_price":5.7,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 304 output tok/s at high thinking; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"model_id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","benchmarks":{"aaii_v4_1":36,"swe_bench_pro":54.2,"terminal_bench_2_1":54},"quality_score":62.5,"quality_coverage":0.55,"quality_vs_fable":65.9,"task_speed_index":null,"speed_score":null,"api_input":0.3,"api_cached_input":0.03,"api_output":2.5,"blended_price":1.84,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 490 output tok/s; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"}],"visualizations":{"bubble":{"x":"cost_score","y":"speed_score","size":"quality_vs_selected_reference","color":"provider","include_when":"quality vs reference, speed score and cost score are all non-null"},"heatmap":{"rows":"model","columns":["quality_vs_selected_reference","speed_score","cost_score"],"color_scale":"sequential_good","domain":[0,100]},"bcg":{"x":"capability_compound","y":"mean(quality_vs_selected_reference, speed_score, cost_score)","color":"provider","quadrants":"medians of currently visible models"}},"capability":{"axes":["coding","reasoning_architecture","knowledge_research","communication_docs","multimodal","agentic"],"focus_multiplier":2.5,"models":[{"model_id":"fable-5","scores":[98,99,96,96,90,99],"capability_compound":96.3,"scq_compound":36.2},{"model_id":"gpt-5.6-sol","scores":[98,97,95,95,96,98],"capability_compound":96.5,"scq_compound":55.7},{"model_id":"gpt-5.6-terra","scores":[95,92,91,92,90,95],"capability_compound":92.5,"scq_compound":70.8},{"model_id":"gpt-5.6-luna","scores":[92,88,86,89,86,91],"capability_compound":88.7,"scq_compound":92.6},{"model_id":"opus-4.8","scores":[92,94,94,95,82,94],"capability_compound":91.8,"scq_compound":54},{"model_id":"gemini-3.1-pro","scores":[78,84,94,84,98,72],"capability_compound":85,"scq_compound":null}],"provenance_note":"Capability scores are editorial rubric ratings synthesized from benchmark and feature evidence, not vendor benchmark results. Keep them separate from benchmark values."},"economics":{"effort_factors":{"none":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4},"cache_hit_presets":{"cold":0,"mixed":0.4,"warm":0.7,"hot":0.9},"models":[{"model_id":"fable-5","base_blended_price":38,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"],"note":"Always-on adaptive thinking; effort is a behavioral control, not a published token multiplier."},{"model_id":"gpt-5.6-sol","base_blended_price":22.5,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-terra","base_blended_price":9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-luna","base_blended_price":0.9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"opus-4.8","base_blended_price":19,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"]},{"model_id":"kimi-k3","base_blended_price":11.4,"cache_read_ratio":0.1,"efforts":["max"],"note":"Always-on reasoning at launch; the max-only effort control is behavioral and not a published token multiplier."}],"efficiency_filters":{"quality_floor_default":75,"provider_default":"all","effort_default":"medium","workload_default":"mixed","sort_default":"scq_compound_desc"}},"hardware_options":[{"id":"local-64gb","name":"64 GB unified-memory workstation","memory_gb":64,"type":"local","best_for":"3B-35B dense or small MoE models at Q4-Q8","cost_note":"Capex varies; compare measured memory bandwidth, not product year."},{"id":"local-128gb","name":"128 GB unified-memory workstation","memory_gb":128,"type":"local","best_for":"Up to roughly 100 GB quantized weights with headroom for KV cache","cost_note":"Portable and quiet; slower than datacenter GPUs for sustained batches."},{"id":"cloud-2xa6000","name":"2 x RTX A6000 48 GB","memory_gb":96,"type":"cloud","best_for":"70B-class dense and mid-size MoE Q4 deployments","cost_note":"Spot price varies by host; record price and interconnect at test time."},{"id":"cloud-h200","name":"1 x H200 141 GB","memory_gb":141,"type":"cloud","best_for":"High-throughput 70B inference and larger quantized MoE models","cost_note":"Use provider quote; hourly rates change frequently."},{"id":"cloud-b200","name":"1 x B200 192 GB","memory_gb":192,"type":"cloud","best_for":"Large-model throughput where software stack supports Blackwell","cost_note":"Availability and hourly rates vary; validate framework support."}],"hardware_model_fit":[{"model_id":"qwen3.6-35b-a3b","hardware_id":"local-64gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":55,"evidence":"planning_estimate"},{"model_id":"qwen3.6-35b-a3b","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"BF16","estimated_output_tps":140,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"local-64gb","fit_status":"unsupported","reason":"Quantized weights exceed usable memory."},{"model_id":"kimi-k2.6-1t","hardware_id":"local-128gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":16,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"cloud-h200","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":36,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-flash-open","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":45,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-h200","fit_status":"unsupported","reason":"Single-device memory is insufficient at the tracked quantization."},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-b200","fit_status":"supported","quant":"Q2_K","estimated_output_tps":14,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"local-64gb","fit_status":"supported","quant":"BF16","estimated_output_tps":38,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"cloud-h200","fit_status":"supported","quant":"BF16","estimated_output_tps":80,"evidence":"planning_estimate"}],"decision_support":{"recommended_stack":[{"route":"hardest_long_horizon","primary":"fable-5","effort":"high","fallback":"gpt-5.6-sol","why":"Use Fable where long-horizon reliability warrants cost and retention policy is acceptable."},{"route":"default_engineering","primary":"gpt-5.6-terra","effort":"medium","fallback":"gpt-5.6-sol","why":"Best balanced default in the current benchmark/cost set."},{"route":"high_volume_subagents","primary":"gpt-5.6-luna","effort":"low","fallback":"gemini-3.5-flash-lite","why":"Luna retains stronger coding-agent evidence; Flash-Lite is the throughput-first fallback for extraction and delegated subtasks."},{"route":"privacy_local","primary":"qwen3.6-35b-a3b","effort":null,"fallback":"deepseek-v4-flash-open","why":"Local-first route where external retention is unacceptable."},{"route":"vision_long_context","primary":"gemini-3.1-pro","effort":"high","fallback":"gpt-5.6-sol","why":"Prefer for multimodal context; validate preview stability before production."}],"routing_rules":["Route by task risk and evidence, not provider family.","Start bulk work on Luna or Terra and escalate only after an explicit verification failure.","Do not send ZDR-required data to Fable 5 because the model requires 30-day retention.","Record model, effort, cache state, region and harness version for every internal speed comparison.","Use self-hosted routes only when the selected hardware row is supported; never infer performance from an empty cell."]},"actions":[{"priority":"P0","owner":"Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","status":"open","order":1},{"priority":"P0","owner":"Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","status":"open","order":2},{"priority":"P0","owner":"Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","status":"open","order":3},{"priority":"P1","owner":"Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","status":"open","order":4},{"priority":"P1","owner":"Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","status":"open","order":5},{"priority":"P2","owner":"Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics.","status":"open","order":6}],"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the price-change history using OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. Added the paid Codex and ChatGPT Work credit reduction, unchanged subscription prices and quota budgets, and Sol API Fast mode (up to 2.5× Standard speed at 2× price)."},{"date":"2026-08-01","tag":"pricing","text":"Refreshed OpenAI's live GPT-5.6 API rate card: Terra is now $2/$12 per million input/output tokens and Luna is $0.20/$1.20, with cached input at $0.20 and $0.02 respectively. Recomputed the 30/70 workload blend, cost-efficiency scores, SCQ compounds and burn baselines. Superseded on August 7 with the official July 30 effective date and full change details."},{"date":"2026-08-01","tag":"benchmark","text":"Recorded DeepSeek V4-Flash-0731's July 31 public-beta evidence: Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in DeepSeek Harness minimal mode. Additional vendor-reported NL2Repo, Cybergym, Toolathlon and Automation Bench results remain narrative evidence rather than new shared columns."},{"date":"2026-07-22","tag":"method","text":"Marked the open-weight quality history as a Fable-anchored source series that is rebased in the browser against the selected reference; fixed task-speed remains explicitly Fable-indexed."},{"date":"2026-07-22","tag":"model","text":"Added Gemini 3.6 Flash and Gemini 3.5 Flash-Lite benchmark, pricing, context and throughput evidence from Google and Artificial Analysis; retained missing task-time fields as null."},{"date":"2026-07-21","tag":"model","text":"Added official Kimi K3 API pricing and derived the documented 30/70 workload blend and cost-efficiency score; no task-speed compound is shown because comparable speed evidence remains unavailable."},{"date":"2026-07-19","text":"Added a dated GPU-hosting price tracker with billing-basis normalization, historical Runpod price changes, current Contabo configurations, and quote-aware Infomaniak options. Replaced stale Vast.ai constants with live-marketplace status."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, context, benchmark records and reference eligibility."},{"date":"2026-07-18","tag":"method","text":"Introduced documented quality, speed, cost, SCQ and six-axis capability composites with explicit missing-data rules."},{"date":"2026-07-18","tag":"hardware","text":"Replaced blank hardware fit speeds with supported/unsupported states and evidence-labelled planning estimates."},{"date":"2026-07-18","tag":"routing","text":"Updated recommended routing to Fable for hardest retained-data work, Terra for default engineering and Luna for high-volume subagents."}],"sources":[{"id":"openai-gpt-5-6","title":"GPT-5.6 launch and evaluations","url":"https://openai.com/index/gpt-5-6/","accessed":"2026-07-18"},{"id":"moonshot-kimi-k3-pricing","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing","accessed":"2026-07-21"},{"id":"openai-models","title":"OpenAI model catalog","url":"https://developers.openai.com/api/docs/models","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-terra","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-luna","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-price-performance","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","accessed":"2026-08-07"},{"id":"openai-gpt-5-6-chat-access","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","accessed":"2026-08-07"},{"id":"deepseek-v4-flash-0731","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","accessed":"2026-08-01"},{"id":"deepseek-v4-pricing","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/","accessed":"2026-08-01"},{"id":"anthropic-models","title":"Claude models overview","url":"https://platform.claude.com/docs/en/about-claude/models/overview","accessed":"2026-07-18"},{"id":"anthropic-retention","title":"Claude API and data retention","url":"https://platform.claude.com/docs/en/manage-claude/api-and-data-retention","accessed":"2026-07-18"},{"id":"google-gemini-3-5","title":"Gemini 3.5 Flash","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash","accessed":"2026-07-18"},{"id":"google-gemini-3-6","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash","accessed":"2026-07-22"},{"id":"google-gemini-3-5-flash-lite","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite","accessed":"2026-07-22"},{"id":"google-flash-july-2026","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/","accessed":"2026-07-22"},{"id":"aa-gemini-3-6-flash","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash","accessed":"2026-07-22"},{"id":"aa-gemini-3-5-flash-lite","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","accessed":"2026-07-22"},{"id":"moonshot-kimi-k3","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3","accessed":"2026-07-18"},{"id":"moonshot-models","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models","accessed":"2026-07-18"},{"id":"openai-gpt-5-5","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/","accessed":"2026-07-18"}],"hosting_prices":{"as_of":"2026-07-19","hours_per_month":730,"methodology":"Preserve provider currency and billing basis. normalized_hourly is a comparison-only derivation for fixed monthly plans; monthly_equivalent multiplies hourly rates by 730. It excludes tax, storage, egress, public IPs, support, discounts, utilization, and model throughput. Append observations instead of replacing them.","trend_policy":"Track the same provider, GPU, service tier, region basis, and currency. A changed SKU starts a new series. Two or more observations produce a trend; one observation is a dated baseline.","offers":[{"id":"runpod-h100-sxm-secure","provider":"Runpod","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$2.99/hr","normalized_hourly":2.99,"monthly_equivalent":2182.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":3.99,"source_n":64},{"date":"2026-07-19","value":2.99,"source_n":63}]},{"id":"runpod-a100-sxm-secure","provider":"Runpod","configuration":"1 x A100 SXM","gpu":"A100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$1.49/hr","normalized_hourly":1.49,"monthly_equivalent":1087.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.94,"source_n":64},{"date":"2026-07-19","value":1.49,"source_n":63}]},{"id":"runpod-l40s-secure","provider":"Runpod","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$0.99/hr","normalized_hourly":0.99,"monthly_equivalent":722.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.19,"source_n":64},{"date":"2026-07-19","value":0.99,"source_n":63}]},{"id":"hyperstack-h100-sxm","provider":"Hyperstack","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$2.40/hr","normalized_hourly":2.4,"monthly_equivalent":1752,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":2.4,"source_n":68}]},{"id":"hyperstack-h200-sxm","provider":"Hyperstack","configuration":"1 x H200 SXM","gpu":"H200 SXM","gpu_count":1,"vram_gb":141,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$3.50/hr","normalized_hourly":3.5,"monthly_equivalent":2555,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":3.5,"source_n":68}]},{"id":"contabo-l40s-monthly","provider":"Contabo","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€751/mo","normalized_hourly":1.0288,"monthly_equivalent":751,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":1.0288,"source_n":65}]},{"id":"contabo-h100-monthly","provider":"Contabo","configuration":"1 x H100","gpu":"H100","gpu_count":1,"vram_gb":80,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€1,838/mo","normalized_hourly":2.5178,"monthly_equivalent":1838,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.5178,"source_n":65}]},{"id":"contabo-h200-nvl-monthly","provider":"Contabo","configuration":"1 x H200 NVL","gpu":"H200 NVL","gpu_count":1,"vram_gb":141,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€2,149/mo","normalized_hourly":2.9438,"monthly_equivalent":2149,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.9438,"source_n":65}]},{"id":"infomaniak-l40s","provider":"Infomaniak","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Switzerland","billing":"Usage-based Public Cloud","price_basis":"Live calculator / availability validation","currency":"CHF","current_price_label":"Calculator / validate","normalized_hourly":null,"monthly_equivalent":null,"price_status":"not_publicly_exposed","checked":"2026-07-19","source_n":66,"history":[]},{"id":"vast-a6000-market","provider":"Vast.ai","configuration":"1 x RTX A6000","gpu":"RTX A6000","gpu_count":1,"vram_gb":48,"region":"Marketplace","billing":"Per second; on-demand, reserved, or interruptible","price_basis":"Live host marketplace","currency":"USD","current_price_label":"Live marketplace","normalized_hourly":null,"monthly_equivalent":null,"price_status":"dynamic_marketplace","checked":"2026-07-19","source_n":69,"history":[]}]}},"model_roster":{"schema_version":"2.0-preview","researched_at":"2026-08-07","default_reference":"fable-5","inclusion_policy":{"default_limit_per_provider":3,"default_scope":"core","summary":"Show a flagship, a balanced or fast model, and one distinctive specialist per provider. Keep siblings and emerging models in expandable extended and watchlist scopes.","promotion_rule":"Promote a watchlist family when at least two signals hold: current official release, meaningful adoption or discussion, differentiated capability or efficiency, and reproducible availability."},"speed_methodology":{"dimensions":["time_to_first_token","output_tokens_per_second","end_to_end_task_time"],"rule":"Compare speed only at the exact model, configuration, provider, region, date, workload, and cache condition. Vendor claims are labelled and unknown values remain unknown.","bands":{"fast":"Designed or measured for low interaction latency","balanced":"General-purpose latency and quality trade-off","deliberate":"Higher reasoning depth or multi-agent execution","unknown":"No comparable evidence"}},"models":[{"id":"fable-5","name":"Claude Fable 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","reference":true,"context":"1M","price":"$10 / $50","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"deliberate","speed_note":"Anthropic comparative latency: slower. No Fable fast mode.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"30-day retention; no zero-data-retention option","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"opus-5","name":"Claude Opus 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Anthropic comparative latency: moderate. Research-preview fast mode claims up to 2.5x higher output throughput at $10 / $50 per MTok.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"Adaptive thinking is on by default; disabling it is supported only through high effort. Independent benchmark results vary materially by effort level and should not be treated as a single generic score.","source":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5"},{"id":"opus-4.8","name":"Claude Opus 4.8","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Optional API fast mode claims up to 2.5x output throughput at premium pricing.","availability":"GA","source":"https://platform.claude.com/docs/en/build-with-claude/fast-mode"},{"id":"sonnet-5","name":"Claude Sonnet 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$3 / $15","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"fast","speed_note":"Anthropic comparative latency: fast.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"haiku-4.5","name":"Claude Haiku 4.5","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"200K","price":"$1 / $5","license":"Proprietary","control":"fixed","levels":[],"default_level":"n/a","speed_band":"fast","speed_note":"Anthropic’s latest verified Haiku and fastest listed Claude; no Haiku 5 is in the current official catalog.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports max reasoning completes comparable AAII work in 61% less time than Fable 5. API Fast mode, introduced July 30, claims up to 2.5x Standard speed at 2x price with unchanged intelligence; no comparable output-tokens/s figure is published.","availability":"GA; ChatGPT, Codex, API","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$2 / $12; cached input $0.20","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 20% from $2.50/$15; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; subscription prices and quota budgets unchanged","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$0.20 / $1.20; cached input $0.02","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"fast","speed_note":"OpenAI positions Luna as the fastest tier and reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 80% from $1/$6; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; rolling out as Free and Go default with unlimited text chats and a Think option","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.5","name":"GPT-5.5","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M API / 400K Codex","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports GPT-5.4-class per-token latency; Codex fast mode is 1.5x throughput at 2.5x cost.","availability":"API, ChatGPT and Codex","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"extended","status":"stable","context":"1M","price":"$30 / $180","license":"Proprietary","control":"reasoning effort","levels":["medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Higher-accuracy tier for difficult work; no comparable public output-throughput figure.","availability":"API and eligible ChatGPT plans","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"harness_bundled","scope":"extended","status":"stable","context":"400K","price":"Included by plan / routed service","license":"Proprietary","control":"fixed","levels":[],"default_level":"provider default","speed_band":"fast","speed_note":"Low-latency ChatGPT route; do not equate service behavior with the reasoning model API.","availability":"ChatGPT","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"preview","context":"1M","price":"$2 / $12","license":"Proprietary","control":"thinking level","levels":["low","medium","high"],"default_level":"high","speed_band":"deliberate","speed_note":"Google warns high thinking may significantly delay the first answer token.","availability":"Preview","source":"https://ai.google.dev/gemini-api/docs/thinking"},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$1.50 / $7.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Artificial Analysis measured about 304 output tok/s at high thinking; Google reports fewer reasoning turns and tool calls than 3.5 Flash.","availability":"GA via Gemini API, AI Studio, Gemini app and Antigravity","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$1.50 / $9","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Still offered; migrate new general Flash workloads to 3.6 Flash for lower output price and stronger agentic performance.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash"},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$0.30 / $2.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Artificial Analysis measured about 490 output tok/s; optimized for high-volume subagents, document parsing and extraction.","availability":"GA via Gemini API, AI Studio and Gemini app rollout","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$0.25 / $1.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Still offered; Gemini 3.5 Flash-Lite is the current high-throughput migration target.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite"},{"id":"grok-4.5","name":"Grok 4.5","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"500K","price":"$2 / $6","license":"Proprietary","control":"reasoning effort","levels":["low","medium","high"],"default_level":"high","speed_band":"fast","speed_note":"xAI reports 80 output tokens/s; vendor measurement conditions are not fully comparable here.","availability":"API and integrations; verify regional availability","source":"https://x.ai/news/grok-4-5"},{"id":"grok-4.20-multi-agent","name":"Grok 4.20 Multi-Agent","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"extended","status":"specialist","context":"1M","price":"$1.25 / $2.50","license":"Proprietary","control":"agent count","levels":["low","medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Levels control a 4- or 16-agent process, not ordinary reasoning depth.","availability":"API","source":"https://docs.x.ai/developers/model-capabilities/text/multi-agent"},{"id":"composer-2.5","name":"Composer 2.5","family":"Composer","provider":"Cursor","region":"US","channel":"harness_bundled","scope":"core","status":"stable","context":"Managed by Cursor","price":"$0.50 / $2.50 standard","license":"Proprietary","control":"speed variant","levels":["standard","fast"],"default_level":"fast","speed_band":"fast","speed_note":"Fast is the default; Cursor documents no public low/medium/high effort selector.","availability":"Cursor only","source":"https://cursor.com/blog/composer-2-5"},{"id":"kimi-k3","name":"Kimi K3","family":"Kimi K3","provider":"Moonshot AI","region":"China","channel":"commercial_api_open_weights_announced","scope":"core","status":"stable","context":"1M","price":"$3 / $15; cached input $0.30","license":"Terms pending weight release","control":"reasoning effort","levels":["max"],"default_level":"max","speed_band":"deliberate","speed_note":"No comparable K3 output-rate figure at launch; max thinking is always enabled.","availability":"API, Kimi, Kimi Work and Kimi Code; weights promised by July 27","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"id":"kimi-k2.7-code","name":"Kimi K2.7 Code","family":"Kimi K2.7","provider":"Moonshot AI","region":"China","channel":"commercial_api","scope":"core","status":"specialist","context":"256K","price":"Current API pricing","license":"Proprietary API","control":"service variant","levels":["standard","high-speed"],"default_level":"standard","speed_band":"fast","speed_note":"High-speed service is documented at about 180 tok/s and up to 260 tok/s for short contexts.","availability":"Kimi API and Kimi Code","source":"https://platform.kimi.ai/docs/models"},{"id":"deepseek-v4","name":"DeepSeek V4","family":"DeepSeek V4","provider":"DeepSeek","region":"China","channel":"open_weight","scope":"core","status":"public_beta","context":"1M","price":"Flash $0.14 / $0.28; cached input $0.0028","license":"MIT","control":"mode","levels":["non-think","think","max"],"default_level":"think","speed_band":"balanced","speed_note":"V4-Flash-0731 reports 82.7 on Terminal-Bench 2.1 at max effort in DeepSeek Harness minimal mode; comparable latency remains unknown.","availability":"V4-Flash-0731 API public beta and weights; future 2x peak-hours pricing announced without an effective date","variants":["Pro","Flash"],"source":"https://api-docs.deepseek.com/updates/"},{"id":"qwen-3.6","name":"Qwen3.6","family":"Qwen3.6","provider":"Alibaba Qwen","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"mode and token budget","levels":["non-thinking","thinking","budget"],"default_level":"thinking","speed_band":"balanced","speed_note":"Budget is provider-native; do not translate it into invented effort labels.","availability":"API and weights","variants":["27B","35B-A3B"],"source":"https://github.com/QwenLM/Qwen3.6"},{"id":"glm-5.2","name":"GLM-5.2","family":"GLM-5","provider":"Z.ai","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"MIT","control":"effort","levels":["adaptive"],"default_level":"adaptive","speed_band":"unknown","speed_note":"Architecture efficiency claims are not a comparable API latency measurement.","availability":"API and weights","source":"https://z.ai/blog/glm-5.2"},{"id":"kimi-k2.5","name":"Kimi K2.5","family":"Kimi K2","provider":"Moonshot AI","region":"China","channel":"open_weight","scope":"extended","status":"deprecated","context":"256K","price":"API and self-hosted","license":"Modified MIT","control":"mode","levels":["instant","thinking"],"default_level":"thinking","speed_band":"balanced","speed_note":"No comparable official throughput measurement; Instant is the lower-latency mode.","availability":"Existing users only; platform sunset August 31, 2026","source":"https://platform.kimi.ai/docs/models"},{"id":"minimax-m3","name":"MiniMax M3","family":"MiniMax M3","provider":"MiniMax","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"Community; non-commercial by default","control":"thinking mode","levels":["disabled","adaptive","enabled"],"default_level":"adaptive","speed_band":"balanced","speed_note":"Vendor claims decode improvements versus M2, not cross-provider latency.","availability":"API and weights","caveat":"License is not permissive open source","source":"https://www.minimax.io/blog/minimax-m3"},{"id":"mistral-small-4","name":"Mistral Small 4","family":"Mistral 3/4","provider":"Mistral AI","region":"EU","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"reasoning effort","levels":["none","high"],"default_level":"none","speed_band":"fast","speed_note":"Vendor claims 40% lower completion time and 3x requests/s versus Small 3.","availability":"API and weights","source":"https://mistral.ai/it/news/mistral-small-4/"},{"id":"llama-4","name":"Llama 4","family":"Llama 4","provider":"Meta","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"10M advertised for Scout","price":"Self-hosted","license":"Llama 4 Community License","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Performance depends on deployment; no comparable official API latency.","availability":"Weights","variants":["Scout","Maverick"],"caveat":"Not OSI-open; EU multimodal and large-platform restrictions apply","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"id":"gemma-3","name":"Gemma 3","family":"Gemma 3","provider":"Google","region":"US","channel":"open_weight","scope":"extended","status":"stable","context":"128K","price":"Self-hosted","license":"Google Gemma Terms","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Deployment dependent; optimized for edge and local use.","availability":"Weights","variants":["1B","4B","12B","27B"],"source":"https://ai.google.dev/gemma/docs/core/model_card_3"},{"id":"gpt-oss","name":"gpt-oss","family":"gpt-oss","provider":"OpenAI","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"128K","price":"Self-hosted","license":"Apache-2.0 plus usage policy","control":"reasoning effort","levels":["low","medium","high"],"default_level":"medium","speed_band":"unknown","speed_note":"Depends on hardware and serving stack.","availability":"Weights","variants":["120b","20b"],"source":"https://openai.com/index/introducing-gpt-oss/"},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","family":"Apertus v1.1","provider":"Swiss AI Initiative","region":"Switzerland","channel":"open_weight","scope":"core","status":"stable","context":"4K","price":"Self-hosted","license":"Apache-2.0","control":"sampling parameters","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"No comparable throughput measurement is published; performance depends on hardware, runtime and quantization.","availability":"Weights; Transformers, vLLM, SGLang and quantized checkpoints","variants":["0.5B","1.5B","4B"],"source":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"id":"inkling","name":"Inkling","family":"Inkling","provider":"Thinking Machines Lab","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"No metered API; open weights + Tinker fine-tuning platform","license":"Apache 2.0","control":"n/a","levels":[],"default_level":null,"speed_band":"unknown","speed_note":"No published throughput figure; not yet independently benchmarked for this roster.","availability":"Open weights (Hugging Face), Tinker fine-tuning platform, third-party inference providers","source":"https://thinkingmachines.ai/model-card/inkling/"}],"watchlist":[{"name":"Step 3.5 Flash","provider":"StepFun","region":"China","license":"Apache-2.0","signal":"Official 100-300 output tok/s claim","source":"https://github.com/stepfun-ai/Step-3.5-Flash"},{"name":"MiMo-V2-Flash","provider":"Xiaomi","region":"China","license":"Apache-2.0","signal":"Thinking toggle and generation-speed claim","source":"https://github.com/XiaomiMiMo/MiMo-V2-Flash"},{"name":"Seed-OSS-36B","provider":"ByteDance","region":"China","license":"Apache-2.0","signal":"512K context and controllable thinking budget","source":"https://github.com/ByteDance-Seed/seed-oss"},{"name":"Hy3 Preview","provider":"Tencent","region":"China","license":"Tencent Hy Community License","signal":"Preview multimodal MoE family","source":"https://github.com/Tencent-Hunyuan/Hy3-preview"}]}}&lt;/script&gt;
&lt;section class="tab-panel" id="tab-models"&gt;
 &lt;h2 class="sect-head"&gt;Current roster&lt;/h2&gt;
 &lt;p class="sect-sub"&gt;A deliberately compact provider roster: flagship, balanced or fast, and differentiated specialist models. Search, filter, or sort the columns; linked values open the primary provider evidence.&lt;/p&gt;</description></item><item><title>Quick Start</title><link>https://projectious-work.github.io/ai-market-research/docs/quick-start/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/docs/quick-start/</guid><description>&lt;p&gt;Requires &lt;code&gt;python3&lt;/code&gt;, &lt;code&gt;bash&lt;/code&gt;, &lt;code&gt;git&lt;/code&gt;, and (for deploys) the &lt;code&gt;gh&lt;/code&gt; CLI. Building
and serving the Hugo/Docsy site additionally requires &lt;code&gt;hugo&lt;/code&gt; (extended) and
&lt;code&gt;node&lt;/code&gt;/&lt;code&gt;npm&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id="build-the-report"&gt;Build the report&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Build the dashboard from the JSON data inputs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;bash src/scripts/build.sh
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Validate JSON + rebuild + sanity-check the artifact&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;bash src/scripts/release-check.sh
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Open the current report directly, without the documentation site&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;xdg-open dist/dashboard.html
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="build-and-serve-the-documentation-site"&gt;Build and serve the documentation site&lt;/h2&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Serve the Hugo/Docsy site locally at http://localhost:1313/, embedding a&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# freshly built report&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;bash docs/scripts/serve-docs.sh
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Produce a production build under docs/public/&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;bash docs/scripts/build-docs.sh
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;docs/scripts/build-docs.sh&lt;/code&gt; and &lt;code&gt;docs/scripts/serve-docs.sh&lt;/code&gt; both rebuild
&lt;code&gt;dist/dashboard.html&lt;/code&gt; via &lt;code&gt;src/scripts/build.sh&lt;/code&gt; and copy it to
&lt;code&gt;static/report/dashboard.html&lt;/code&gt; before invoking Hugo, so the embedded report in
&lt;a href="https://projectious-work.github.io/ai-market-research/report/"&gt;Signal Room&lt;/a&gt; always reflects the current data.&lt;/p&gt;</description></item><item><title>02 Tools</title><link>https://projectious-work.github.io/ai-market-research/report/02-tools/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/report/02-tools/</guid><description>&lt;div class="sr-report-scope"&gt;
&lt;script id="market-data" type="application/json"&gt;{"meta":{"generated_at":"2026-08-07T00:00:00Z","reference_default":"fable-5","report_metrics_file":"data/report-metrics.json"},"executive_summary":{"models":["**Claude Opus 5 is now generally available.** Anthropic's new `claude-opus-5` targets complex agentic coding and enterprise work with a 1M-token context window, 128K maximum output, adaptive thinking by default, and unchanged $5/$25 per MTok base pricing. Independent benchmark values are not yet recorded here.","**Google reset the Flash price-performance curve on July 21.** Gemini 3.6 Flash is GA at $1.50/$7.50 per million tokens with stronger coding and agentic results than 3.5 Flash, while Gemini 3.5 Flash-Lite reaches roughly 490 output tok/s at $0.30/$2.50 for high-volume subagents and extraction.","**Kimi K3 resets the open-weight frontier.** Moonshot reports a 2.8T sparse MoE, native vision, 1M context, and 99% of Fable 5 across 14 overlapping launch-suite evaluations; weights are promised by July 27.","**GPT-5.6 Luna and Terra became materially cheaper on July 30.** Luna fell 80% to $0.20/$1.20 per MTok and Terra 20% to $2/$12; their paid Codex and ChatGPT Work usage also consumes fewer credits, while subscription prices and quota budgets did not change.","**GPT-5.6 now spans Sol, Terra and Luna**, giving leaders a deliberate capability, balanced, and high-throughput ladder under one family. Sol API Fast mode replaces Priority Processing: OpenAI claims up to 2.5× Standard speed at 2× price, with unchanged intelligence.","**Luna is expanding beyond API routing.** OpenAI says it becomes the default for Free and Go users, with unlimited text chats and a higher-reasoning Think option rolling out subject to abuse guardrails. This is a ChatGPT product update, not an API capability change.","**Keep benchmark tables separate by harness.** OpenAI's GPT-5.6 launch results and Scale's public SWE-bench Pro leaderboard use different model versions and evaluation setups; use each as evidence, not as one directly rankable series.","**GPT-5.5 remains active in Codex and the API.** It is retained as a compatibility and portfolio option rather than being hidden by the newer 5.6 family.","**Claude Haiku 4.5 remains Anthropic’s latest verified Haiku.** No official Haiku 5 listing was found in Anthropic’s current model catalog."],"harnesses":["**Separate model capability from harness capability.** Tool execution, context management, isolation, and observability can dominate real workflow outcomes.","**Subscription access is not a production routing contract.** Validate API, credit-pool, and third-party harness policies before standardizing an operating model.","**Maintain at least one portable fallback path** across providers for high-value workflows and operational incidents.","**Gemini Managed Agents now add background execution, remote MCP, custom functions and credential refresh.** That makes Google a more credible managed-agent control plane, but it does not substitute for workload-specific evaluation."],"self_hosting":["**Open weights are now a strategic option, not only a cost play.** Kimi K3, DeepSeek, Qwen, Nemotron, Mistral, and Llama cover different sovereignty and specialization needs.","**Do not compare self-hosting at zero token cost.** Include accelerator rental or depreciation, power, utilization, serving staff, and measured throughput.","**Pilot against a defined workload and hardware envelope** before treating advertised context or parameter scale as deployable capacity."],"strategy":["**Run a portfolio, not a winner-takes-all model standard.** Reserve frontier reasoning for high-value decisions and route routine work to measured fast or efficient tiers.","**Instrument quality, latency, retries, and total workflow cost together.** Token price alone is not an operating metric.","**Review the portfolio quarterly and after major releases**, with explicit retirement, security, and fallback criteria."]},"headline_stats":[{"id":"frontier_count","label":"Models tracked","value":35,"unit":"","delta":"Current, fast and retained compatibility models","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_weight","label":"Open-weight","color":"#16866f"}],"history":[{"date":"2026-05-18","proprietary":20,"open_weight":3},{"date":"2026-07-22","proprietary":26,"open_weight":8},{"date":"2026-07-25","proprietary":27,"open_weight":8}],"note":"Release-tag roster snapshots; announced open-weight releases are counted with open-weight models."}},{"id":"best_open_pct","label":"Best open-weight vs selected reference","value":99,"unit":"%","delta":"Kimi K3; geometric mean across 14 overlapping vendor evals","delta_dir":"up","signal_id":"open_weight_quality"},{"id":"cheapest_frontier_api","label":"Lowest frontier blended API price","value":0.9,"unit":"$/Mtok","delta":"GPT-5.6 Luna; 30% input / 70% output blend","delta_dir":"down","signal_id":"blended_frontier_price"},{"id":"fastest_task_rate","label":"Fastest task-rate index","value":"3.2×","unit":"","delta":"Fable 5 = 1.0×; planning index, not tokens/second","delta_dir":"up","signal_id":"task_speed"},{"id":"max_context","label":"Largest usable context","value":"10M","unit":"tok","delta":"Llama 4 Scout; deployment constraints still apply","delta_dir":"up","trend":[{"date":"2026-01-01","value":1},{"date":"2026-04-01","value":1},{"date":"2026-07-01","value":10}]},{"id":"harness_count","label":"Harnesses tracked","value":16,"unit":"","delta":"Commercial and open agent environments","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_source","label":"Open source","color":"#3d78c5"}],"history":[{"date":"2026-05-18","proprietary":3,"open_source":11},{"date":"2026-07-22","proprietary":6,"open_source":10}],"note":"Release-tag roster snapshots classified from each harness license."}},{"id":"documented_output_speed","label":"Fastest documented API output","value":490,"unit":"tok/s","delta":"Gemini 3.5 Flash-Lite; Artificial Analysis first-party API measurement","delta_dir":"up","signal_id":"output_throughput"},{"id":"quality_coverage","label":"Models with quality evidence","value":"31/35","unit":"","delta":"Unknown remains unknown; no zero-value substitution","delta_dir":"neutral","trend":[{"date":"2026-01-01","value":18},{"date":"2026-04-01","value":24},{"date":"2026-07-01","value":31}]}],"trends":{"best_open_vs_opus":{"label":"Best open-weight % vs Opus 4.7","history":[{"date":"2025-11-01","value":68},{"date":"2025-12-01","value":72},{"date":"2026-01-01","value":78},{"date":"2026-02-01","value":80},{"date":"2026-03-01","value":83},{"date":"2026-04-01","value":86},{"date":"2026-05-01","value":88},{"date":"2026-06-01","value":90}]},"median_frontier_output_price":{"label":"Median frontier $/Mtok output","history":[{"date":"2025-11-01","value":18},{"date":"2025-12-01","value":17},{"date":"2026-01-01","value":15.5},{"date":"2026-02-01","value":14.5},{"date":"2026-03-01","value":13},{"date":"2026-04-01","value":12.5},{"date":"2026-05-01","value":12},{"date":"2026-06-01","value":11}]},"models_released_per_month":{"label":"Notable model releases per month","history":[{"date":"2025-11-01","value":3},{"date":"2025-12-01","value":4},{"date":"2026-01-01","value":5},{"date":"2026-02-01","value":6},{"date":"2026-03-01","value":4},{"date":"2026-04-01","value":7},{"date":"2026-05-01","value":5},{"date":"2026-06-01","value":8}]}},"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the GPT-5.6 pricing history from OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. The report now records lower paid Codex and ChatGPT Work credit consumption, unchanged subscription prices and quota budgets, Sol API Fast mode (up to 2.5× Standard speed at 2× price), and the August 6 ChatGPT Luna access expansion."},{"date":"2026-08-01","tag":"pricing","text":"Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to $0.20/$1.20 from OpenAI's live model pages. Recomputed the tracked 30/70 workload blends and effort-burn ratios. Superseded on August 7 with OpenAI's published July 30 effective date and full price-change details."},{"date":"2026-08-01","tag":"model","text":"Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output, thinking/non-thinking modes, $0.14/$0.28 per-million pricing, $0.0028 cached input, and vendor-reported agent results including Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2."},{"date":"2026-07-28","tag":"model","text":"Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B active) open-weight (Apache 2.0) multimodal MoE with 1M-token context, released 2026-07-15. No first-party per-token API pricing is published (monetized via the Tinker fine-tuning platform); benchmark scores are published on the model card but not yet normalized into this roster's comparable set, so quality/speed/cost fields are left unknown rather than estimated."},{"date":"2026-07-25","tag":"model","text":"Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive thinking by default, $5/$25 per MTok base pricing, and official cloud-platform availability. Added independent Artificial Analysis evidence (61 Intelligence Index at max effort; 52.3 output tok/s) with effort-specific caveats. Gemini 3.5 Flash Cyber remains limited to CodeMender government and trusted-partner pilots; GPT-Live and Muse Spark 1.1 remain non-API products, so none were added to the API roster. Claude Opus 4.7 Fast Mode was removed July 24."},{"date":"2026-07-24","tag":"benchmark","text":"Refreshed current-source evidence: added OpenAI's GPT-5.6 launch table, Google's Managed Agents update, and Scale's public SWE-bench Pro leaderboard. Clarified that vendor launch tables and the public leaderboard are not directly comparable because their model versions and harnesses differ."},{"date":"2026-07-22","tag":"fix","text":"Made the selected reference propagate through open-weight quality headlines, market-signal history, comparison headings, model analytics, self-hosting quality and capability market position. Replaced the model and harness inventory mini-lines with stacked proprietary/open category areas based on release-tag roster snapshots."},{"date":"2026-07-22","tag":"model","text":"Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M context, 64K output, current API pricing, Artificial Analysis intelligence and throughput measurements, and Google’s published coding and agentic benchmarks."},{"date":"2026-07-22","tag":"data","text":"Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to governments and trusted CodeMender partners; Gemini Omni Flash and Nano Banana 2 Lite remain specialized media models rather than general-purpose roster entries."},{"date":"2026-07-21","tag":"model","text":"Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint: fully open Apache 2.0 weights and data, 4K context, 1.7T-token distillation, 1,811 languages, and official BF16, FP8, NVFP4A16, INT3, INT4 and INT6 variants."},{"date":"2026-07-21","tag":"harness","text":"Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI 0.144.6, Gemini CLI 0.51.0, OpenCode 1.18.4, Cline 4.0.10, and Goose 1.43.0; updated repository star snapshots and notable release capabilities."},{"date":"2026-07-21","tag":"model","text":"Added Moonshot's official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached input, and $15/M output; the 30/70 workload blend is $11.40/M before reasoning-effort effects."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark registers and selectable reference configurations. Added documented quality, speed, cost and capability composites in data/report-metrics.json."},{"date":"2026-07-18","tag":"routing","text":"Updated the action queue and recommended routing: Fable for hardest retained-data workloads, Terra for default engineering, Luna for high-volume subagents, with explicit escalation rules."},{"date":"2026-06-06","tag":"fix","text":"Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K context, 1000+ tok/s on Cerebras) and Hermes Agent v0.16.0 (Nous Research, MIT, self-hosted multi-platform agent) — both were incorrectly removed in v1.5.0 sweep."},{"date":"2026-06-06","tag":"policy","text":"Dashboard market sweep v1.5.0: real-world re-grounding. Replaced fictional Mythos/GPT-5.5-Cyber rows with verified models; added Nvidia Nemotron coalition, Kimi K2.6, GLM-5, Cohere Command A+, SubQ 1M-Preview."},{"date":"2026-06-04","tag":"model","text":"Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context, NVIDIA Open Model License) at Computex — first frontier-scale open model from Nvidia."},{"date":"2026-06-04","tag":"model","text":"Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral, Perplexity, Reflection AI, Sarvam, Thinking Machines Lab as inaugural members."},{"date":"2026-06-01","tag":"model","text":"Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model."}],"actions":["P0 · Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","P0 · Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","P0 · Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","P1 · Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","P1 · Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","P2 · Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics."],"models":[{"id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","tier":"frontier","released":"2026-06-09","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":80,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii":59.9,"coding_agent_index":77.2,"deep_swe":69.7,"terminal_bench":83.1,"agents_last_exam":40.5,"api_in":10,"api_out":50,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":52.3,"subscription":"Claude API and supported cloud platforms","notes":"Default reference for v2. Anthropic describes Fable 5 as its most capable widely released model. Adaptive thinking is always on. Comparative latency is slower. Retention is 30 days and zero-data-retention is not available.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"status":"stable","speed_class":"deliberate","speed_evidence":"vendor-qualitative","capability_levels":{"coding":49,"reasoning":50,"knowledge":48,"comms":48,"multimodal":45,"agentic":50}},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":64.6,"swe_verified":null,"aaii":58.9,"coding_agent_index":80,"deep_swe":72.7,"terminal_bench":88.8,"agents_last_exam":52.7,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Flagship GPT-5.6 tier. API Fast mode replaced Priority Processing on July 30: OpenAI claims up to 2.5× Standard speed at 2× Standard price with no intelligence change. Speed is otherwise stored as an end-to-end task index in report-metrics.json, not tok/s.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":49,"knowledge":48,"comms":48,"multimodal":48,"agentic":49}},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":63.4,"swe_verified":null,"aaii":55,"coding_agent_index":77.4,"deep_swe":69.6,"terminal_bench":87.4,"agents_last_exam":50.4,"api_in":2,"api_out":12,"api_cache_hit":0.2,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Balanced GPT-5.6 tier and recommended default engineering route. On July 30, API list pricing fell 20% from $2.50/$15 to $2/$12 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":48,"reasoning":46,"knowledge":46,"comms":46,"multimodal":45,"agentic":48}},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","tier":"fast","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":62.7,"swe_verified":null,"aaii":51.2,"coding_agent_index":74.6,"deep_swe":67.2,"terminal_bench":84.7,"agents_last_exam":50.3,"api_in":0.2,"api_out":1.2,"api_cache_hit":0.02,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Fastest, most affordable GPT-5.6 tier; recommended for high-volume subagents with verification and escalation. On July 30, API list pricing fell 80% from $1/$6 to $0.20/$1.20 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change. ChatGPT is also rolling Luna out as the Free and Go default, with unlimited text chats and a Think option; that product change does not alter API routing.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":46,"reasoning":44,"knowledge":43,"comms":45,"multimodal":43,"agentic":46}},{"id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","tier":"frontier","released":"2026-07-16","license":"open weights announced; terms pending weight release","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":null,"tok_per_sec":null,"subscription":"Kimi API, Kimi Code, Kimi Work","reasoning_capable":true,"effort_default":"max","effort_levels":["max"],"notes":"2.8T sparse MoE with 16/896 experts active, native vision and always-on thinking. API pricing is $3/M uncached input, $0.30/M cached input and $15/M output. Full weights promised by July 27, 2026.","quality_vs_fable":99.2,"quality_evidence":"Geometric mean across 14 overlapping values in Moonshot’s launch comparison; vendor-reported, max/xhigh settings.","capability_levels":{"coding":47,"reasoning":46,"knowledge":47,"comms":46,"multimodal":47,"agentic":48}},{"id":"opus-5","name":"Claude Opus 5","provider":"Anthropic","tier":"frontier","released":"2026-07-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"Claude API, Bedrock, Google Cloud, Microsoft Foundry","notes":"Current Anthropic Opus generation. 1M context and 128K maximum output. Adaptive thinking is enabled by default; effort defaults to high. Artificial Analysis reports a 61 Intelligence Index and 52.3 output tok/s at max effort; these figures are effort-specific. Research-preview Fast Mode is Claude API-only at $10/$50 per MTok and claims up to 2.5x higher output throughput.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"speed_class":"balanced","speed_evidence":"vendor-qualitative","cite":[86,87]},{"id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","tier":"frontier","released":"2026-05-28","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":69.2,"swe_verified":88.6,"livecodebench":82,"aime":90,"tau2":86,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Max 20× $200/mo · Max 5× $100/mo","notes":"Previous Opus generation, retained as an active compatibility option. SWE-V 88.6%, SWE-Pro 69.2%, AAII 61.4. Fast Mode reduced 3× to $10/$50 (was $30/$150 on 4.7). 1M context standard.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":50,"comms":50,"multimodal":34,"agentic":50}},{"id":"opus-4.7","name":"Claude Opus 4.7","provider":"Anthropic","tier":"frontier","released":"2026-04-16","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":64.3,"swe_verified":87.6,"livecodebench":79,"aime":88,"tau2":84,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":50,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Now legacy as of Opus 4.8 release May 28. Pricing unchanged. SWE-Verified 87.6%, SWE-Pro 64.3%. Fast Mode was removed July 24, 2026; standard-speed API access remains active.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":49,"comms":50,"multimodal":32,"agentic":50}},{"id":"sonnet-4.6","name":"Claude Sonnet 4.6","provider":"Anthropic","tier":"frontier","released":"2026-02-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":44,"swe_verified":77,"livecodebench":76,"aime":85,"tau2":81,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":50,"tok_per_sec":80,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Best code style/intent understanding. With cache+batch: $0.30/$7.50 effective.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","max"],"cache_discount":0.1,"capability_levels":{"coding":41,"reasoning":40,"knowledge":39,"comms":42,"multimodal":31,"agentic":41}},{"id":"haiku-4.5","name":"Claude Haiku 4.5","provider":"Anthropic","tier":"fast","released":"2025-10-15","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":28,"swe_verified":62,"livecodebench":58,"aime":70,"tau2":65,"api_in":1,"api_out":5,"api_cache_hit":0.1,"batch_discount":50,"tok_per_sec":110,"subscription":"Available in all tiers","notes":"5× cheaper than Sonnet. Triage/classification champion.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.1,"capability_levels":{"coding":28,"reasoning":18,"knowledge":29,"comms":30,"multimodal":18,"agentic":28}},{"id":"gpt-5.5","name":"GPT-5.5","provider":"OpenAI","tier":"frontier","released":"2026-04-23","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M (272K threshold)","swe_pro":58.6,"swe_verified":88.7,"livecodebench":84,"aime":92,"tau2":82,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Pro $200 · Pro Lite $100 · Plus $20","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"OpenAI flagship; default in ChatGPT (Instant variant since May 5, 2026). 1M context with 2× input/1.5× output surcharge above 272K. Reasoning tokens billed as output.","cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":41,"knowledge":49,"comms":40,"multimodal":41,"agentic":38}},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","provider":"OpenAI","tier":"frontier","released":"2026-04-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":60,"swe_verified":90,"livecodebench":86,"aime":94,"tau2":84,"api_in":30,"api_out":180,"api_cache_hit":3,"batch_discount":50,"tok_per_sec":35,"subscription":"Pro $200 only","reasoning_capable":true,"effort_default":"high","effort_levels":["medium","high","xhigh"],"notes":"Highest-stakes reasoning tier. $30/$180. Available in ChatGPT Pro $200 and as API. 6× cost of base 5.5.","cache_discount":0.25,"capability_levels":{"coding":47,"reasoning":48,"knowledge":50,"comms":42,"multimodal":42,"agentic":42}},{"id":"gpt-5.4","name":"GPT-5.4","provider":"OpenAI","tier":"frontier","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":57.7,"swe_verified":81,"livecodebench":81,"aime":90,"tau2":80,"api_in":2.5,"api_out":15,"api_cache_hit":0.25,"batch_discount":50,"tok_per_sec":60,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Best quality-per-credit. Held SWE-Pro lead Feb-April. Default for everyday coding.","cache_discount":0.1,"capability_levels":{"coding":40,"reasoning":40,"knowledge":40,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-mini","name":"GPT-5.4 Mini","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":38,"swe_verified":73,"livecodebench":71,"aime":78,"tau2":70,"api_in":0.4,"api_out":1.6,"api_cache_hit":0.04,"batch_discount":50,"tok_per_sec":130,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"94% of GPT-5.4 coding at 6× less. Best subagent. ~1/20 burn vs GPT-5.5. Collapses at 64K+ context.","cache_discount":0.1,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-nano","name":"GPT-5.4 Nano","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":25,"swe_verified":60,"livecodebench":58,"aime":65,"tau2":55,"api_in":0.1,"api_out":0.4,"api_cache_hit":0.01,"batch_discount":50,"tok_per_sec":200,"subscription":"API only","reasoning_capable":true,"effort_default":"low","effort_levels":["minimal","low","medium"],"notes":"Smallest reasoning model. API-only. For embeddable/edge inference at near-zero cost.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":20,"comms":20,"multimodal":20,"agentic":20}},{"id":"gpt-5.3-codex","name":"GPT-5.3-Codex","provider":"OpenAI","tier":"frontier","released":"2026-01-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":56.8,"swe_verified":85,"livecodebench":82,"aime":87,"tau2":78,"api_in":1.5,"api_out":10,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":70,"subscription":"All paid tiers · Code Review uses this","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Coding-specialised. ~⅓ burn vs GPT-5.5 for ~2pts less SWE-Pro. Often best $/quality for execution turns.","cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":28,"knowledge":28,"comms":27,"multimodal":8,"agentic":38}},{"id":"gpt-5.3-codex-spark","name":"GPT-5.3-Codex-Spark","provider":"OpenAI","tier":"fast","released":"2026-02-12","license":"proprietary","jurisdiction":"US","context":128000,"context_label":"128K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":1000,"subscription":"ChatGPT Pro — research preview only","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"Smaller, latency-first sibling of GPT-5.3-Codex. Released Feb 12, 2026. 1000+ tok/s on Cerebras hardware. ChatGPT Pro research preview only — not in API at launch; separate preview rate-limit pool (no standard credit burn). Text-only. Target use: real-time micro-edits, live pair-programming in Codex app/CLI/VS Code.","cache_discount":null,"capability_levels":{"coding":38,"reasoning":22,"knowledge":22,"comms":25,"multimodal":0,"agentic":28}},{"id":"gpt-5.2","name":"GPT-5.2","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":52,"swe_verified":78,"livecodebench":76,"aime":84,"tau2":75,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Available but not recommended","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Codex picker keeps it for 'long-running agents' (specifically tuned for autonomy). Otherwise eclipsed by 5.3-Codex.","cache_discount":0.1},{"id":"gpt-5.2-codex","name":"GPT-5.2-Codex","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":50,"swe_verified":76,"livecodebench":75,"aime":82,"tau2":73,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Legacy","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Predecessor to 5.3-Codex. Legacy.","cache_discount":0.1},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","provider":"OpenAI","tier":"fast","released":"2026-05-05","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":35,"swe_verified":76,"livecodebench":72,"aime":81,"tau2":70,"api_in":1.5,"api_out":6,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":200,"subscription":"Default ChatGPT model · API chat-latest","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"ChatGPT default since May 5, 2026. 52.5% fewer hallucinations vs 5.3, 30% shorter responses, first Instant-class High-capability rating on cybersec/bio-chem.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":40,"comms":30,"multimodal":30,"agentic":10}},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","provider":"Google","tier":"frontier","released":"2026-02-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":54.2,"swe_verified":78,"livecodebench":76,"aime":85,"tau2":78,"api_in":2,"api_out":12,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":119,"subscription":"Gemini Advanced $20 · Ultra $100","notes":"Released Feb 19, 2026 (preview). Pricing doubles above 200K input tokens. 50% batch discount.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":39,"reasoning":41,"knowledge":49,"comms":41,"multimodal":50,"agentic":29}},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":58.7,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78,"agents_last_exam":null,"quality_vs_fable":81.6,"api_in":1.5,"api_out":7.5,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":304,"subscription":"Gemini API · AI Studio · Gemini app · Antigravity","notes":"GA stable ID gemini-3.6-flash. Google reports fewer tool calls and 17% fewer output tokens than 3.5 Flash on the AA Index workload; Computer Use is preview. Artificial Analysis measured about 304 output tok/s at high thinking.","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":42,"reasoning":38,"knowledge":45,"comms":40,"multimodal":47,"agentic":42}},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","provider":"Google","tier":"legacy","released":"2026-05-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55.1,"swe_verified":78,"livecodebench":70,"aime":75,"tau2":68,"aaii_v4_1":50,"swe_bench_pro":55.1,"deep_swe_v1_1":37,"terminal_bench_2_1":76.2,"agents_last_exam":null,"quality_vs_fable":76.1,"api_in":1.5,"api_out":9,"api_cache_hit":0.375,"batch_discount":50,"tok_per_sec":165,"subscription":"Free CLI: 1000 req/day at 1M context","notes":"Shipped GA at Google I/O May 19, 2026. Still offered, but Gemini 3.6 Flash is the recommended migration target with stronger agentic results and lower output-token pricing.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":30,"reasoning":30,"knowledge":40,"comms":30,"multimodal":40,"agentic":20}},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":54.2,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":36,"swe_bench_pro":54.2,"deep_swe_v1_1":null,"terminal_bench_2_1":54,"agents_last_exam":null,"quality_vs_fable":65.9,"api_in":0.3,"api_out":2.5,"api_cache_hit":0.03,"batch_discount":50,"tok_per_sec":490,"subscription":"Gemini API · AI Studio · Gemini app rollout","notes":"GA stable ID gemini-3.5-flash-lite. Google positions it for high-volume subagents, document parsing and structured extraction; Artificial Analysis measured about 490 output tok/s. Computer Use availability differs across Google documentation and should be validated per API surface.","reasoning_capable":true,"effort_default":"minimal","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":32,"reasoning":28,"knowledge":34,"comms":32,"multimodal":38,"agentic":35}},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash Lite","provider":"Google","tier":"legacy","released":"2026-03-10","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":22,"swe_verified":65,"livecodebench":60,"aime":68,"tau2":58,"api_in":0.25,"api_out":1.5,"api_cache_hit":0.0625,"batch_discount":50,"tok_per_sec":220,"subscription":"Free CLI + AI Studio","notes":"Lite tier; 1M context retained.","reasoning_capable":true,"effort_default":"off","effort_levels":["off","low","medium"],"cache_discount":0.25,"capability_levels":{"coding":22,"reasoning":22,"knowledge":30,"comms":25,"multimodal":35,"agentic":15}},{"id":"deepseek-v4-pro","name":"DeepSeek V4-Pro","provider":"DeepSeek","tier":"frontier","released":"2026-04-24","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":55.4,"swe_verified":80.6,"livecodebench":93.5,"aime":88,"tau2":78,"api_in":0.435,"api_out":0.87,"api_cache_hit":0.043,"batch_discount":null,"tok_per_sec":60,"subscription":"API only","notes":"MIT-licensed. Permanent pricing May 22, 2026 — Opus-class quality at ~1/10 cost. 1.6T/49B MoE, 1M context.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.0083,"capability_levels":{"coding":47,"reasoning":40,"knowledge":38,"comms":30,"multimodal":8,"agentic":28}},{"id":"deepseek-v4-flash","name":"DeepSeek V4-Flash","provider":"DeepSeek","tier":"fast","released":"2026-07-31","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"tau2":72,"terminal_bench_2_1":82.7,"agents_last_exam":25.2,"api_in":0.14,"api_out":0.28,"api_cache_hit":0.0028,"batch_discount":null,"tok_per_sec":90,"subscription":"API + open weights","notes":"V4-Flash-0731 public beta; 284B/13B active MoE, 1M context and 384K max output. DeepSeek reports Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in its unreleased minimal harness; internal DSBench results are not normalized here. A future 2x peak-hours rate is announced, but no effective date is published.","reasoning_capable":true,"effort_default":"think","effort_levels":["non-think","think"],"cache_discount":0.02,"capability_levels":{"coding":40,"reasoning":30,"knowledge":30,"comms":25,"multimodal":5,"agentic":25}},{"id":"minimax-m2.7","name":"MiniMax M2.7","provider":"MiniMax","tier":"frontier","released":"2026-03-18","license":"open-weight","jurisdiction":"China","context":205000,"context_label":"205K","swe_pro":40,"swe_verified":74,"livecodebench":72,"aime":80,"tau2":73,"api_in":0.3,"api_out":1.2,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":70,"subscription":"API + open-weight","notes":"Current flagship reasoner; MoE 230B/10B active.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.52,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":20,"agentic":20}},{"id":"grok-4.3","name":"Grok 4.3","provider":"xAI","tier":"frontier","released":"2026-05-06","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":45,"swe_verified":80,"livecodebench":80,"aime":88,"tau2":76,"api_in":1.25,"api_out":2.5,"api_cache_hit":0.31,"batch_discount":null,"tok_per_sec":75,"subscription":"SuperGrok $30 · Heavy $300","notes":"Current xAI flagship — aggressive pricing for frontier tier. Hybrid reasoning. Multimodal text+image.","reasoning_capable":true,"effort_default":"reasoning-on","effort_levels":["off","on"],"cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":44,"knowledge":42,"comms":34,"multimodal":34,"agentic":34}},{"id":"grok-4.1-fast","name":"Grok 4.1 Fast","provider":"xAI","tier":"fast","released":"2026-03-20","license":"proprietary","jurisdiction":"US","context":2000000,"context_label":"2M","swe_pro":30,"swe_verified":70,"livecodebench":65,"aime":75,"tau2":65,"api_in":0.2,"api_out":0.5,"api_cache_hit":0.05,"batch_discount":null,"tok_per_sec":140,"subscription":"SuperGrok $30","notes":"Cheapest large-context model on market. 2M context, $0.05/Mtok cached input.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.25,"capability_levels":{"coding":28,"reasoning":28,"knowledge":30,"comms":26,"multimodal":26,"agentic":22}},{"id":"mistral-medium-3.5","name":"Mistral Medium 3.5","provider":"Mistral","tier":"frontier","released":"2026-04-29","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":42,"swe_verified":77.6,"livecodebench":75,"aime":75,"tau2":65,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":75,"subscription":"Le Chat Pro €20","notes":"EU jurisdiction. 128B dense, Apache 2.0. Strongest non-Chinese open-weight coding agent. Vibe agents (GitHub/Linear/Jira/Sentry integrations). Pricing not verified; Medium 3 was $0.40/$2.00.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":40,"reasoning":30,"knowledge":40,"comms":30,"multimodal":10,"agentic":20}},{"id":"mistral-large-3","name":"Mistral Large 3","provider":"Mistral","tier":"frontier","released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"tau2":70,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":65,"subscription":"API + Le Chat","notes":"Cheapest premium output price in market among Western frontier. 675B/41B MoE, Apache 2.0.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":38,"reasoning":36,"knowledge":40,"comms":36,"multimodal":20,"agentic":28}},{"id":"subq-1m-preview","name":"SubQ 1M-Preview","provider":"SubQ","tier":"frontier","released":"2026-05-15","license":"proprietary","jurisdiction":"US","context":12000000,"context_label":"12M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":1,"api_out":5,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":80,"subscription":"API preview","notes":"First commercial subquadratic (non-transformer) LLM. ~1/5 frontier cost on long-context tasks. Capability rating estimated — public benchmarks pending.","reasoning_capable":null,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":30,"reasoning":32,"knowledge":36,"comms":30,"multimodal":5,"agentic":28}},{"id":"nemotron-3-ultra","name":"Nvidia Nemotron 3 Ultra","provider":"Nvidia","tier":"frontier","released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"tau2":76,"api_in":1.5,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":300,"subscription":"build.nvidia.com (closed API tier) + open weights","notes":"550B/55B MoE, hybrid Mamba-Transformer. ~300 tok/s. Open weights also available — see self_hosting. Capability levels are best estimates pending independent benchmarks.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":42,"reasoning":42,"knowledge":42,"comms":34,"multimodal":10,"agentic":34}},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","provider":"Swiss AI Initiative","tier":"fast","released":"2026-06-15","license":"Apache 2.0","jurisdiction":"Switzerland","context":4096,"context_label":"4K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Self-hosted weights","notes":"Largest newly released Apertus v1.1 distilled checkpoint. Dense 4.6B storage / 3.8B compute parameters, trained on 1.7T tokens, supports 1,811 languages, and ships in BF16 plus server and Apple-oriented quantizations. No comparable coding-agent, throughput, or API-price evidence is published.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null},{"id":"inkling","name":"Inkling","provider":"Thinking Machines Lab","tier":"frontier","released":"2026-07-15","license":"Apache 2.0","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Open weights (Hugging Face); fine-tuning and inference via the Tinker platform and third-party providers","notes":"975B-parameter multimodal MoE (41B active), 66-layer decoder-only transformer, 1M-token context. Text/image/audio input, text-only output. No first-party per-token API pricing published -- Thinking Machines monetizes via the Tinker fine-tuning platform rather than metered inference. A smaller Inkling-Small (12B active) companion model was released alongside it. Benchmark results are published on the model card across reasoning, agentic, coding, factuality, vision, audio, and safety categories but are not yet normalized into this roster's comparable benchmark set.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null}],"subscriptions":[{"provider":"Anthropic","tier":"Pro","price_usd":20,"limits":"~45 messages / 5h on Sonnet · limited Opus","models":"Sonnet 4.6, Haiku 4.5, limited Opus 4.8/4.7","features":"Chat only (Claude Code removed April 2026). From June 15, 2026 split into Chat pool + Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 5×","price_usd":100,"limits":"5× Pro quotas · ~225 msg/5h Sonnet · expanded Opus","models":"Full Opus 4.8/4.7 · Sonnet 4.6 · Haiku 4.5","features":"Cache reads included flat-rate. From June 15, 2026: Chat pool + separate Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 20×","price_usd":200,"limits":"20× Pro quotas","models":"All","features":"For heavy Opus users. From June 15, 2026: Chat + Agent SDK credit pools."},{"provider":"OpenAI","tier":"Plus","price_usd":20,"limits":"~80 GPT-5.4 msg/3h","models":"GPT-5.4 (limited) · GPT-5.4 Mini · o-series","features":"ChatGPT · GPTs · Codex CLI 30-150 tasks/5h"},{"provider":"OpenAI","tier":"Pro 5×","price_usd":100,"limits":"5× Plus quotas (new tier April 2026)","models":"GPT-5.4 Thinking unlimited","features":"Released as Anthropic Max competitor"},{"provider":"OpenAI","tier":"Pro","price_usd":200,"limits":"Effectively unlimited","models":"All including o3-pro","features":"Original premium tier"},{"provider":"Google","tier":"Gemini Advanced","price_usd":20,"limits":"Generous, soft caps","models":"Gemini 3.1 Pro · Gemini 3.6 Flash · Gemini 3.5 Flash-Lite","features":"Workspace integration · 1M context"},{"provider":"Google","tier":"Ultra","price_usd":100,"limits":"Higher quotas + Veo video","models":"All Gemini + research preview","features":"Veo 3 video · Project Mariner"},{"provider":"Mistral","tier":"Le Chat Pro","price_usd":22,"limits":"Generous","models":"Mistral Large 3 · Codestral","features":"EU jurisdiction"},{"provider":"xAI","tier":"SuperGrok","price_usd":30,"limits":"Generous","models":"Grok 4 · Grok 4 Heavy","features":"X integration"},{"provider":"DeepSeek","tier":"API only","price_usd":null,"limits":"Pay per token","models":"V4-Pro · V4-Flash","features":"Cheapest frontier API ($0.435/$0.87 permanent since May 22, 2026)"}],"agent_policies":[{"provider":"Anthropic","subscription_automated":"Prohibited","enforcement":"Active (OAuth blocked Apr 4)","first_party_exception":"`claude -p` pipe mode and Claude Code itself","api_required_for_automation":true,"cite":[35]},{"provider":"OpenAI","subscription_automated":"Prohibited (ToS)","enforcement":"Currently tolerated","first_party_exception":"Codex CLI uses subscription quota","api_required_for_automation":false},{"provider":"Google","subscription_automated":"Allowed in CLI","enforcement":"—","first_party_exception":"Gemini CLI free tier 1000 req/day","api_required_for_automation":false},{"provider":"Mistral","subscription_automated":"Allowed","enforcement":"—","first_party_exception":"—","api_required_for_automation":false}],"harnesses":[{"id":"claude-code","name":"Claude Code","vendor":"Anthropic","license":"proprietary","category":"CLI + IDE","stars":49000,"providers":["Anthropic"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":true,"computer_use":true,"lsp":true,"git":true,"memory":"CLAUDE.md","sandbox":"local","swe_pro":46,"pricing":"Subscription Pro/Max","sweet_spot":"SubagentStop hooks, /plugin list, requiredMinimumVersion managed setting, MCP fixes, Agent Teams.","stumbles":"Anthropic-only. Removed from standard Pro tier April 2026 — push to Max.","cite":[23,1]},{"id":"opencode","name":"OpenCode","vendor":"Anomaly","license":"MIT","category":"CLI + ACP","stars":188119,"providers":["75+: Anthropic*, OpenAI, Google, Mistral, Kimi, GLM, Ollama, LM Studio"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.18.4 (Jul 20, 2026): adaptive Kimi thinking controls, provider-defined reasoning options, restored Azure endpoints, and a rewritten desktop prompt input.","stumbles":"Anthropic OAuth blocked April 4 — must use API key.","cite":[24,35,72]},{"id":"codex-cli","name":"Codex CLI","vendor":"OpenAI","license":"Apache 2.0","category":"CLI + macOS app","stars":100232,"providers":["OpenAI"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"cloud","swe_pro":56.8,"pricing":"Plus $20 / Pro $200","sweet_spot":"v0.144.6 (Jul 18, 2026): refreshed GPT-5.6 Sol/Terra/Luna bundled instructions and corrected Codex context-window metadata to 272K.","stumbles":"OpenAI-only. No MCP, no hooks. Tightly coupled to apply_patch tool.","cite":[25,36,70]},{"id":"gemini-cli","name":"Gemini CLI","vendor":"Google","license":"Apache 2.0","category":"CLI","stars":106096,"providers":["Google"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"GEMINI.md","sandbox":"local","swe_pro":null,"pricing":"Free 1000 req/day","sweet_spot":"v0.51.0 (Jul 16, 2026): hardened sensitive-path and symlink handling, read-only macOS sandbox git config, and modern-model escape-sequence fixes.","stumbles":"Sunsetting to Antigravity CLI for free tier on June 18, 2026; paid Gemini/Enterprise keys retain access.","cite":[26,71]},{"id":"aider","name":"Aider","vendor":"paul-gauthier","license":"Apache 2.0","category":"CLI","stars":32000,"providers":["Anthropic*","OpenAI","Google","Ollama","100+"],"mcp":false,"skills":false,"hooks":false,"subagents":false,"voice":true,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"CONVENTIONS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"Mature, lightweight, voice-input native. Pair-programmer mode. Active, last commit Mar 2026. 44k stars.","stumbles":"No MCP, no hooks. Less ambitious than newer harnesses.","cite":[27]},{"id":"cline","name":"Cline","vendor":"cline-bot","license":"Apache 2.0","category":"VS Code extension","stars":64886,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".clinerules","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v4.0.10 (Jul 20, 2026): current release adds telemetry for consecutive-mistake-limit events; multi-editor and CLI surfaces remain available.","stumbles":"Anthropic OAuth blocked. Can be expensive on long sessions.","cite":[28,35,73]},{"id":"roo-code","name":"Roo Code","vendor":"RooVetGit","license":"Apache 2.0","category":"VS Code extension","stars":18000,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".roo","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v3.53.0 (Apr 23, 2026). Power-user Cline fork, model-agnostic, BYOK. Recent: GPT-5.5 via Codex, Opus 4.7 on Vertex, checkpoint nav.","stumbles":"Same OAuth situation. Configuration complexity.","cite":[29,35]},{"id":"cursor","name":"Cursor","vendor":"Anysphere","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".cursorrules","sandbox":"local","swe_pro":null,"pricing":"Hobby free · Pro $20 · Pro+ $60 · Ultra $200 · Teams $40/user","sweet_spot":"Cursor 3.5 (May 20, 2026): Cloud Agents (isolated VMs, multi-repo), Composer 2.5, Agents Window, parallel subagents.","stumbles":"Closed source. Lock-in. Subscription required for serious use.","cite":[31]},{"id":"windsurf","name":"Windsurf / Devin Desktop","vendor":"Cognition","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"windsurfrules","sandbox":"local","swe_pro":null,"pricing":"Pro $20 · Max $200","sweet_spot":"Renamed to Devin Desktop, Agent Command Center kanban, embedded Devin cloud agent, SWE-1.6 model, multi-agent + worktrees.","stumbles":"Pro $20 (was $15); smaller ecosystem than Cursor.","cite":[32]},{"id":"goose","name":"Goose","vendor":"Block","license":"Apache 2.0","category":"Desktop + CLI","stars":51387,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":false,"git":true,"memory":".goosehints","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.43.0 (Jul 14, 2026): per-message token/cost/TTFT/tok-s metrics, ACP reconnection, GPT-5.6 support, dynamic Ollama Cloud discovery, and expanded providers.","stumbles":"Anthropic OAuth blocked. Less mindshare than OpenCode.","cite":[30,35,74]},{"id":"omo","name":"OMO (Multi-model orchestrator)","vendor":"community","license":"MIT","category":"Multi-agent orchestrator","stars":54000,"providers":["multi-provider"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free","sweet_spot":"Rebrand from oh-my-opencode; multi-model orchestration. Wraps Claude Code, OpenCode, Codex, Kimi K2, DeepSeek V4, Gemini CLI.","stumbles":"Niche. Steep learning curve."},{"id":"hermes","name":"Hermes Agent","vendor":"Nous Research","license":"MIT","category":"Self-hosted multi-platform agent","stars":null,"providers":["OpenRouter-style multi-model"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":true,"remote":true,"computer_use":true,"lsp":false,"git":true,"memory":"persistent memory + auto-gen skills","sandbox":"local/docker/ssh/singularity/modal","swe_pro":null,"pricing":"Free (self-hosted)","sweet_spot":"v0.16.0. Persistent memory + auto-generated skills — learns your projects. Bridges Telegram/Discord/Slack/WhatsApp/Signal/Email/CLI. Natural-language cron for unattended runs. Parallel isolated subagents. Web search, browser automation, vision, image-gen, TTS.","stumbles":"Not coding-specialised — general autonomous agent. Setup requires self-host script. No first-party SWE benchmark.","cite":[]},{"id":"github-copilot-cli","name":"GitHub Copilot CLI","vendor":"GitHub","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["GitHub Models"],"mcp":false,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":true,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Bundled with Copilot Business/Enterprise","sweet_spot":"GA Feb 25, 2026. Specialized sub-agents (Explore, Task, Code Review, Plan), background delegation, autopilot.","stumbles":"Bundled-only — no standalone tier. GitHub-centric."},{"id":"amp-cli","name":"Amp CLI","vendor":"Sourcegraph","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["Multi"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Subscription (Sourcegraph)","sweet_spot":"Spun out as standalone company 2026. Runs as sidebar agent inside Zed via Terminal Threads.","stumbles":"Early standalone phase."},{"id":"zed","name":"Zed","vendor":"Zed Industries","license":"proprietary","category":"IDE","stars":null,"providers":["15 LLM providers + MCP"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Personal free (2k predictions) · Pro $10 · Business $30/seat","sweet_spot":"Rust-native editor with first-class agent panel + ACP host. Terminal Threads run Claude Code/Amp inline.","stumbles":"Editor first; agent layer still maturing."},{"id":"continue-dev","name":"Continue","vendor":"Continue","license":"Apache 2.0","category":"IDE extension + CLI","stars":null,"providers":["multi-provider"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".continue","sandbox":"local","swe_pro":null,"pricing":"Solo $0 · Team/Company ~$10/dev/mo","sweet_spot":"Agent mode plan+execute. Continuous AI, Mission Control, shared PR/ticket workflows.","stumbles":"Newer agent features still stabilizing across providers."}],"self_hosting":{"hardware_options":[{"id":"vast-2xa6000","name":"Vast.ai 2× A6000","vram_gb":96,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Reference cloud-GPU setup for this dashboard. No upfront capex. Docker templates, SSH/Cloudflare Zero Trust. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-3xa6000","name":"Vast.ai 3× A6000","vram_gb":144,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Headroom for Qwen 3 235B-A22B Q6_K. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-4xa6000","name":"Vast.ai 4× A6000","vram_gb":192,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Llama 3.1 405B Q4 territory. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"mbp-14-m4pro-64","name":"MacBook Pro 14\" M4 Pro 64GB","vram_gb":64,"cost_label":"~$3,200 capex","cost_per_hour":null,"type":"local","notes":"MoE sweet spot. 273 GB/s memory bandwidth."},{"id":"mbp-16-m5max-128","name":"MacBook Pro 16\" M5 Max 128GB","vram_gb":128,"cost_label":"~$5,500 capex","cost_per_hour":null,"type":"local","notes":"Best portable inference. ~545 GB/s bandwidth."},{"id":"mba-15-32","name":"MacBook Air 15\" 32GB","vram_gb":32,"cost_label":"~$1,900 capex","cost_per_hour":null,"type":"local","notes":"Hard 32GB ceiling. Limited to ~14B dense or 26B MoE Q4."},{"id":"contabo-2xl40s","name":"Contabo 2 x L40S","provider":"Contabo","vram_gb":96,"cost_label":"€1,502/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"EU-oriented fixed monthly configuration; 64 vCPU, 213 GB RAM, 3.5 TB storage, and 15 TB bandwidth are listed for the 2-GPU tier."},{"id":"contabo-1xh200","name":"Contabo 1 x H200 NVL","provider":"Contabo","vram_gb":141,"cost_label":"€2,149/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"Fixed monthly large-memory option. Confirm location and availability before treating it as a sovereignty or latency fit."},{"id":"infomaniak-1xl40s","name":"Infomaniak 1 x L40S","provider":"Infomaniak","vram_gb":48,"cost_label":"Live calculator / availability validation","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"calculator_or_request","price_checked":"2026-07-19","source":"https://www.infomaniak.com/en/hosting/public-cloud/prices","notes":"Swiss OpenStack option with dedicated GPU access and usage billing. Public pages list L40S availability but do not expose a stable crawlable SKU price; validate availability for the selected region."},{"id":"hyperstack-1xh200","name":"Hyperstack 1 x H200 SXM","provider":"Hyperstack","vram_gb":141,"cost_label":"$3.50/hr on demand","cost_per_hour":3.5,"type":"cloud","show_in_fit":false,"price_status":"published_on_demand","price_checked":"2026-07-19","source":"https://www.hyperstack.cloud/","notes":"Minute-accurate on-demand billing; reservation pricing starts at $2.45/hr. Validate region, storage, and availability."}],"models":[{"id":"gemma-4-26b-moe","name":"Gemma 4 26B-A4B MoE","params_total":26,"params_active":4,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":35,"livecodebench":77,"aime":88,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-3xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-4xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":22,"tok_per_sec":75},"mbp-16-m5max-128":{"quant":"BF16","vram_used":52,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"MoE — only 4B active per token. Faster than dense models 5× its size."},{"id":"gemma-4-31b-dense","name":"Gemma 4 31B Dense","params_total":31,"params_active":31,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":38,"livecodebench":80,"aime":89,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-3xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-4xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":19,"tok_per_sec":28},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":26,"tok_per_sec":38},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Highest quality open Gemma. Slower per-token (full 31B active)."},{"id":"qwen-3.6-plus","name":"Qwen 3.6 Plus","params_total":397,"params_active":17,"released":"2026-04-11","license":"Apache 2.0","jurisdiction":"China","swe_pro":50,"livecodebench":71.4,"aime":87,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":38},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":35},"vast-4xa6000":{"quant":"Q8_0","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1M context. Top open agentic coder. April 11 release.","swe_verified":68.2},{"id":"llama-4-maverick","name":"Llama 4 Maverick","params_total":400,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":42,"livecodebench":70,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":28},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":25},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"400B / 17B active MoE. 1M context. Strong MMLU-Pro (80.5%) but coding behind Chinese labs.","swe_verified":72},{"id":"minimax-m2.5-open","name":"MiniMax M2.5 (open weights)","params_total":456,"params_active":46,"released":"2026-01-20","license":"open-weight","jurisdiction":"China","swe_pro":40,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Eclipsed by DeepSeek V4 and Qwen 3.6 Plus. China jurisdiction. Maintained for niche workloads."},{"id":"deepseek-v4-pro-open","name":"DeepSeek V4-Pro (open weights)","params_total":1600,"params_active":49,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":55.4,"livecodebench":93.5,"aime":90,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q2_K","vram_used":188,"tok_per_sec":14},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1.6T total / 49B active MoE. MIT license. Strongest open coder. Requires very large multi-GPU deployment."},{"id":"llama-4-scout","name":"Llama 4 Scout","params_total":109,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":36,"swe_verified":68,"livecodebench":66,"aime":78,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-3xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-4xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":32,"tok_per_sec":22},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":45,"tok_per_sec":32},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"10M token context (longest in any open model). 109B / 17B active MoE."},{"id":"kimi-k2.6","name":"Kimi K2.6","params_total":235,"params_active":21,"released":"2026-05-01","license":"Modified MIT","jurisdiction":"China","swe_pro":47,"swe_verified":75,"livecodebench":78,"aime":84,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":36},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":32},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Best open-weight for sub-agent fan-out. Built for harness-driven parallel pipelines. Chinese-trained."},{"id":"glm-5.1","name":"GLM 5.1","params_total":358,"params_active":32,"released":"2026-04-22","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":76,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT license — rare among open frontier models besides DeepSeek. Strong for enterprise fine-tuning."},{"id":"mistral-small-4","name":"Mistral Small 4","params_total":24,"params_active":24,"released":"2026-04-18","license":"Apache 2.0","jurisdiction":"EU","swe_pro":28,"swe_verified":60,"livecodebench":58,"aime":68,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-3xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-4xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":48,"tok_per_sec":38},"mbp-16-m5max-128":{"quant":"BF16","vram_used":48,"tok_per_sec":55},"mba-15-32":{"quant":"Q4_K_M","vram_used":14,"tok_per_sec":30}},"notes":"6.5B effective parameters. EU jurisdiction. Best on-device option."},{"id":"nemotron-3-ultra-550b-a55b-moe","name":"Nvidia Nemotron 3 Ultra","params_total":550,"params_active":55,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q3_K_M","vram_used":180,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Hybrid Mamba-Transformer. 1M context. NVIDIA Open Model License. Computex June 4 launch."},{"id":"nemotron-3-nano-30b-a3b","name":"Nvidia Nemotron 3 Nano 30B-A3B","params_total":30,"params_active":3,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":30,"swe_verified":70,"livecodebench":68,"aime":76,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":24,"tok_per_sec":70},"mbp-16-m5max-128":{"quant":"BF16","vram_used":60,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":17,"tok_per_sec":50}},"notes":"Hybrid Mamba-Transformer Nano variant. Open weights."},{"id":"nemotron-nano-9b-v2","name":"Nvidia Nemotron Nano 9B v2","params_total":9,"params_active":9,"released":"2026-04-12","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":22,"swe_verified":60,"livecodebench":55,"aime":65,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-3xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-4xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":18,"tok_per_sec":90},"mbp-16-m5max-128":{"quant":"BF16","vram_used":18,"tok_per_sec":120},"mba-15-32":{"quant":"Q4_K_M","vram_used":6,"tok_per_sec":65}},"notes":"Dense 9B. Strong instruction following at edge sizes."},{"id":"kimi-k2.6-1t","name":"Kimi K2.6 (1T MoE)","params_total":1000,"params_active":32,"released":"2026-04-20","license":"Modified MIT","jurisdiction":"China","swe_pro":58.6,"swe_verified":80.2,"livecodebench":82,"aime":88,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Top open intelligence. 80.2% SWE-Verified, 58.6% SWE-Pro, AAII 54. 262K context. Modified MIT."},{"id":"glm-4.6","name":"GLM 4.6","params_total":355,"params_active":32,"released":"2025-09-15","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":75,"aime":80,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":28},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":24},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 200K context. Z.ai release."},{"id":"qwen3.6-35b-a3b","name":"Qwen 3.6 35B-A3B","params_total":35,"params_active":3,"released":"2026-04-16","license":"Apache 2.0","jurisdiction":"China","swe_pro":38,"swe_verified":74,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":22,"tok_per_sec":80},"mbp-16-m5max-128":{"quant":"BF16","vram_used":70,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"Apache 2.0 MoE — strong $/quality for fast bulk inference."},{"id":"cohere-command-a-plus","name":"Cohere Command A+","params_total":218,"params_active":28,"released":"2026-05-22","license":"CC-BY-NC + commercial","jurisdiction":"Canada","swe_pro":42,"swe_verified":75,"livecodebench":70,"aime":78,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":30},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":26},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"First open-weights Cohere release in &gt;1 year. Sparse-MoE multimodal."},{"id":"mistral-large-3-open","name":"Mistral Large 3 (open weights)","params_total":675,"params_active":41,"released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q3_K_M","vram_used":140,"tok_per_sec":22},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":180,"tok_per_sec":20},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"675B/41B active MoE. Apache 2.0. EU jurisdiction."},{"id":"deepseek-v4-flash-open","name":"DeepSeek V4-Flash (open weights)","params_total":284,"params_active":13,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":45},"vast-3xa6000":{"quant":"Q6_K","vram_used":115,"tok_per_sec":40},"vast-4xa6000":{"quant":"Q8_0","vram_used":150,"tok_per_sec":34},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":20},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 1M context, 384K max output. Cheap large-context open option."}],"frameworks":[{"name":"llama.cpp","best_for":"Mac (MLX), broad GGUF support","notes":"Best Apple Silicon performance via Metal."},{"name":"vLLM","best_for":"Multi-GPU servers, throughput","notes":"Production serving. Tensor parallelism."},{"name":"Ollama","best_for":"Easiest setup, dev workflow","notes":"Wrapper around llama.cpp. One-line model pull."},{"name":"MLX","best_for":"Apple Silicon native","notes":"Apple's framework. Best M-series perf."}]},"strategy":{"current_recommendation":{"label":"Risk-tiered GPT-5.6 + Fable stack","monthly_usd":null,"components":["GPT-5.6 Luna low — high-volume subagents and routine transformations","Gemini 3.5 Flash-Lite minimal — throughput-first extraction, parsing and parallel subagents","Gemini 3.6 Flash medium — Google-first coding, computer-use and multimodal agent loops","GPT-5.6 Terra medium — default engineering, review and documentation","GPT-5.6 Sol high or Claude Fable 5 high — escalation for hard, long-horizon tasks","Qwen 3.6 35B-A3B or DeepSeek V4-Flash — local route for privacy-sensitive bulk work"],"rationale":"Route by task risk instead of one subscription. Terra and Luna now retain strong coding-agent quality at substantially lower list price; Fable remains the long-horizon option only where its 30-day retention requirement is acceptable. Monthly cost is workload-dependent and must be computed from measured token volume."},"alternatives":[{"label":"Dual subscription (Claude Max + OpenAI Pro)","monthly_usd":230,"rationale":"Adds GPT-5.4 SWE-bench Pro lead and Codex CLI cloud sandbox. Worth $100/mo only if you frequently hit hard issues where Sonnet 4.6 plateaus.","verdict":"Defer until you have measured Sonnet plateau frequency for 30 days."},{"label":"API-only (no subscriptions)","monthly_usd":200,"rationale":"Pure pay-per-use. Maximum flexibility. Loses the subscription cache advantage — same workload costs 1.5–5× more for power users.","verdict":"Worse economics for your usage volume. Skip."},{"label":"Self-hosted maximalist","monthly_usd":80,"rationale":"Vast.ai 24/7 with Gemma 4 + Qwen 3 + occasional API top-up for frontier-only tasks.","verdict":"Cheapest if quality plateau at ~88% of Opus is acceptable. Operational overhead is real."}],"routing":[{"tier":"Bulk (70%)","use_for":"Classification, simple edits, triage, log parsing","preferred":"Gemini 3.5 Flash-Lite minimal · GPT-5.6 Luna low · Qwen 3.6 35B-A3B when local","cost_label":"Gemini $0.30/$2.50; Luna $0.20/$1.20 Mtok before cache, batch and reasoning tokens"},{"tier":"Mid (25%)","use_for":"Multi-file edits, code review, refactors, docs","preferred":"GPT-5.6 Terra medium · Gemini 3.6 Flash medium for Google-first or multimodal work","cost_label":"Terra $2/$12; Gemini $1.50/$7.50 Mtok before cache, batch and reasoning tokens"},{"tier":"Premium (5%)","use_for":"Architecture, hard debugging, long-context refactors","preferred":"GPT-5.6 Sol high · Claude Fable 5 high for long-horizon autonomy","cost_label":"Sol $5/$30; Fable $10/$50 Mtok before reasoning tokens"}],"open_questions":["Will the Nvidia Nemotron Coalition (Mistral, Cursor, Black Forest Labs, Thinking Machines) actually deliver a frontier-open consortium, or fragment within a quarter?","How does the June 15 Anthropic billing restructure (Chat pool + Agent SDK credit pool) change Max economics for agentic workloads?","Does Gemini 3.6 Flash’s lower output price and stronger agentic performance displace 3.1 Pro for most coding workflows before Gemini 3.5 Pro arrives?","DeepSeek V4-Pro permanent pricing ($0.435/$0.87) — does it force closed providers (OpenAI, Anthropic) to cut headline prices?","SubQ 1M-Preview's subquadratic architecture — does the next year see other commercial non-transformer LLMs?"],"task_fit":{"description":"For each task type, one recommended pick per provider drawn from the full roster. Click any cell to see runners-up that the daily briefing also considered. Burn × is derived from the matrix and recomputes when you change the reference dropdown.","rows":[{"task":"Tiny edits / boilerplate","description":"Single-file syntax fixes, regex replacements, formatting, docstrings, import sorting","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fixed-rate, fastest paid Anthropic option"},"openai":{"model_id":"gpt-5.5","effort":"minimal","rationale":"Cheapest reasoning setting on flagship"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Fastest Google route for cheap, high-volume pattern work"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Effectively $0/tok on existing Vast.ai infra"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ minimal","gpt-5.4-nano @ minimal"],"anthropic":["sonnet-4.6 @ low"],"google":["gemini-3.1-pro @ off"],"self_hosted":["llama-3.1-70b @ Q6_K"]}},{"task":"Normal coding task","description":"Single-function implementation, straightforward bug fixes, simple feature additions","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best $/quality on Anthropic side; cache reads included in Max 5×"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-specialised, ~⅓ burn of GPT-5.5 for near-identical SWE-Pro"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"58.7% SWE-Pro and stronger agentic coding than 3.5 Flash at lower output price"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Highest-quality open Gemma at full BF16 in 96GB"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium","gpt-5.5 @ low"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ high"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Multi-file implementation","description":"Feature spanning 3-10 files, requires understanding cross-file context","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Best style/intent understanding for code spanning files"},"openai":{"model_id":"gpt-5.5","effort":"medium","rationale":"Strong on planning; medium gives consistent multi-file edits"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Improved multi-step coding loops, fewer unwanted edits and 1M context"},"self_hosted":{"model_id":"qwen-3.6-plus","effort":null,"rationale":"Largest open MoE; strong reasoning across files"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ high","gpt-5.4 @ high"],"anthropic":["opus-4.7 @ high"],"google":["gemini-3.1-pro @ high"],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Hard debugging / architecture","description":"Race conditions, performance bottlenecks, design decisions with long-term implications","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Default Claude Code effort for Opus; best long-context retention"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"SWE-Pro leader; high effort balances depth and burn"},"google":{"model_id":"gemini-3.1-pro","effort":"high","rationale":"Preview Pro remains the deepest Google reasoning route; compare 3.6 Flash before paying the premium"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open-weight models trail frontier ~10-15pts here; not yet ready for hardest cases"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh","gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ high (lower burn)"],"google":[],"self_hosted":["qwen-3-235b-a22b — viable for some cases"]}},{"task":"Very hard autonomous repo task","description":"Overnight runs, agentic loops, tasks you don't supervise turn-by-turn","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"max","rationale":"Autonomy earns max effort's premium; no human in the loop to correct"},"openai":{"model_id":"gpt-5.5","effort":"xhigh","rationale":"Highest available reasoning budget on OpenAI side"},"google":{"model_id":null,"effort":null,"rationale":"Gemini 3.6 Flash improves agent loops but still lacks a max-equivalent tier for the hardest unsupervised runs"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Not recommended — frontier models earn their cost on the hardest tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ xhigh — better but $100/mo gating"],"anthropic":["opus-4.7 @ xhigh — if max budget is too steep"],"google":[],"self_hosted":[]}},{"task":"Code review / PR comments","description":"Reviewing diffs, suggesting improvements, finding issues without writing code yourself","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Critical thinking at moderate burn; doesn't need Opus depth"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-tuned reviewer at ⅓ burn of GPT-5.5"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Better instruction following and fewer unwanted code edits than 3.5 Flash"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Critique tasks don't need frontier; quality plateau acceptable"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ medium"],"self_hosted":["qwen-3-235b-a22b"]}},{"task":"Long-context refactor (50K+ tokens)","description":"Cross-cutting changes across a large codebase, where retention is the bottleneck","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"high","rationale":"1M context with retention that actually works; Opus's sweet spot"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"400K API context; degrades faster than Opus past 200K"},"google":{"model_id":"gemini-3.6-flash","effort":"high","rationale":"Google reports a 54% 1M-context MRCR score, roughly double 3.5 Flash"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open models cap around 200K useful context; not yet competitive"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh — if context fits in 256K"],"anthropic":["opus-4.7 @ xhigh"],"google":["gemini-3.1-pro @ high"],"self_hosted":[]}},{"task":"Subagent / delegated subtask","description":"Spawned by a main agent for focused work; quality threshold lower than user-facing","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fast, cheap, no effort knob to manage"},"openai":{"model_id":"gpt-5.4-mini","effort":"medium","rationale":"Best subagent — 94% of GPT-5.4 coding at 6× less"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Purpose-built for high-volume subagents at roughly 490 output tok/s"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Zero per-token cost for high-volume subagent loops"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ low","gpt-5.4-nano @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["llama-3.1-70b @ Q4_K_M"]}},{"task":"Documentation / explanation","description":"READMEs, code comments, tutorials, explaining existing code in prose","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best prose style on Anthropic side"},"openai":{"model_id":"gpt-5.4","effort":"medium","rationale":"Reasoning helps less here; medium-effort GPT-5.4 wins on $/quality"},"google":{"model_id":"gemini-3.6-flash","effort":"low","rationale":"Lower output price than 3.5 Flash with improved knowledge-work and document analysis"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Open models are competitive for prose tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ low"],"anthropic":["haiku-4.5"],"google":["gemini-3.5-flash-lite @ low"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Test scaffolding / fixtures","description":"Boilerplate test files, fixture generation, mocking, parameterised test suites","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Pattern-following work; doesn't need reasoning depth"},"openai":{"model_id":"gpt-5.4-mini","effort":"low","rationale":"Pattern-heavy; minimal reasoning, fast turnaround"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Lowest-latency current Google model for repetitive structured generation"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Bulk test scaffolding is exactly where self-hosted earns out"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Cybersecurity / vulnerability research","description":"Penetration testing, red-teaming, vulnerability discovery — tasks where cyber-specific safeguards apply","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Most capable Anthropic model; cyber tasks require Cyber Verification Program"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"Flagship reasoning; vetted security access via OpenAI"},"google":{"model_id":null,"effort":null,"rationale":"No cyber-specialized Gemini variant"},"self_hosted":{"model_id":"deepseek-v4-pro","effort":null,"rationale":"Open-weight without safety filtering; deploy in isolated environment"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ xhigh"],"google":[],"self_hosted":["qwen-3.6-plus","kimi-k2.6"]}}]}},"quota_burn_matrix":{"baseline_label":"selected reference model at medium effort = 1.00×","baseline_model_id":"fable-5","baseline_effort":"medium","methodology":"Burn ratios are displayed as multiples of the selected reference model at medium effort. Underlying values stored as absolute units anchored to gpt-5.5 medium = 1.00. Per-model ratios from API list pricing (firm, ±5%). Per-effort multipliers grounded in: Anthropic's published thinking-token budgets (low=skip, medium=~1k, high=5k, xhigh=10k, max=20k tokens); ArtificialAnalysis's measurement of Sonnet 4.6 max ≈ 3× Sonnet 4.5 on the Intelligence Index; nxcode.io / OpenAI guidance that xhigh ≈ 3-5× medium; ampcode's GPT-5.5 cost analysis. Per-effort multipliers are ±20%. NOTE: quality vs effort is not strictly monotonic per task — aggregate quality trends upward, but individual tasks can see high beat xhigh or medium beat high. OpenAI explicitly warns that 'high is not automatically better than medium'. Models that run in a separate preview quota bucket with no standard-pool burn are encoded as 0.00x and annotated until final token or credit rates are published.","stacking_multipliers":[{"name":"Fast mode (/fast on)","multiplier":2,"scope":"any cell"},{"name":"Cached input (long thread)","multiplier":0.6,"scope":"any cell"},{"name":"Plan mode (/plan)","multiplier":"uses high effort","scope":"overrides default"}],"openai_matrix":[{"model_id":"gpt-5.6-sol","minimal":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4,"note":"List-price blend equals the historical GPT-5.5 medium unit. Effort factors are planning estimates until measured token use is available."},{"model_id":"gpt-5.6-terra","minimal":0.18,"low":0.26,"medium":0.4,"high":0.68,"xhigh":1.08,"max":1.6,"note":"Two-fifths of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.6-luna","minimal":0.018,"low":0.026,"medium":0.04,"high":0.068,"xhigh":0.108,"max":0.16,"note":"Four percent of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.5","minimal":0.2,"low":0.5,"medium":1,"high":2,"xhigh":3.5},{"model_id":"gpt-5.5-pro","minimal":null,"low":null,"medium":6,"high":12,"xhigh":21},{"model_id":"gpt-5.4","minimal":0.1,"low":0.25,"medium":0.5,"high":1,"xhigh":1.75},{"model_id":"gpt-5.4-mini","minimal":0.011,"low":0.027,"medium":0.053,"high":0.107,"xhigh":0.187},{"model_id":"gpt-5.4-nano","minimal":0.003,"low":0.007,"medium":0.013,"high":null,"xhigh":null},{"model_id":"gpt-5.3-codex","minimal":0.07,"low":0.17,"medium":0.33,"high":0.67,"xhigh":1.17},{"model_id":"gpt-5.2","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null},{"model_id":"gpt-5.2-codex","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null}],"anthropic_matrix":[{"model_id":"fable-5","low":1.1,"medium":1.69,"high":2.87,"xhigh":4.56,"max":6.76,"note":"Always-on adaptive thinking. Values start from the $10/$50 list-price blend; effort multipliers are planning estimates rather than published token budgets."},{"model_id":"opus-4.8","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Current flagship. Same headline $5/$25. Fast Mode 3× cheaper ($10/$50) vs Opus 4.7."},{"model_id":"opus-4.7","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Legacy as of May 28. Thinking tokens: low=skip, medium=~1k, high=5k, xhigh=10k, max=20k. New tokenizer +35% tokens vs 4.6."},{"model_id":"sonnet-4.6","low":0.25,"medium":0.5,"high":1.25,"xhigh":null,"max":3.5,"note":"API default high. No xhigh tier. AA measured max ≈ 3× Sonnet 4.5 cost."},{"model_id":"haiku-4.5","low":null,"medium":0.17,"high":null,"xhigh":null,"max":null,"note":"No effort control — fixed-rate model."}],"google_matrix":[{"model_id":"gemini-3.1-pro","off":0.13,"low":0.2,"medium":0.3,"high":0.5},{"model_id":"gemini-3.5-flash","off":0.02,"low":0.03,"medium":0.05,"high":0.09},{"model_id":"gemini-3.6-flash","off":0.017,"low":0.025,"medium":0.042,"high":0.076},{"model_id":"gemini-3.5-flash-lite","off":0.005,"low":0.008,"medium":0.014,"high":0.024}],"non_reasoning_note":"Models without reasoning controls: Haiku 4.5 (Anthropic matrix above as fixed rate), GPT-5.5 Instant (ChatGPT default since May 5), Mistral Medium 3.5, MiniMax M2.7, Kimi K2.6, GLM 4.6, Qwen 3.6 variants, all Llama 4 variants, Grok 4.1 Fast, SubQ 1M-Preview — burn at fixed rate per model regardless of effort knob. DeepSeek V4-Flash-0731 now exposes thinking and non-thinking modes.","effort_quality_factors":{"minimal":0.65,"low":0.85,"medium":0.94,"high":0.98,"xhigh":1,"max":1.02,"off":0.85,"on":0.95},"quality_methodology":"Quality % = (model_swe_pro / reference_swe_pro × effort_quality_factor) × 100. Effort quality factors: minimal=0.65, low=0.85, medium=0.94, high=0.98, xhigh=1.00, max=1.02. Anchors: Anthropic's Hex measurement (low Opus 4.7 ≈ medium Opus 4.6 quality) anchors low ≈ 0.85; apiyi.com's report that max gains ~3pts over xhigh on hardest tasks anchors max=1.02; ampcode's GPT-5.5 analysis (medium captures 'most' of capability) anchors medium=0.94. IMPORTANT: these are AGGREGATE estimates. stet.sh found per-task reversals — high can beat xhigh on some tasks, medium can beat high. Treat as ±10% indicators, not precise measures.","unit_anchor":"gpt-5.5 medium","unit_anchor_note":"All burn values are stored as absolute units anchored to gpt-5.5 medium = 1.00. At render time, each cell is divided by the reference model's medium-effort cell (or its single datapoint for non-reasoning models) to produce the displayed ×-of-reference number. When the reference dropdown changes, all burn ratios recalculate against the new reference.","workload_presets":{"cold":{"label":"Cold (0% cache)","description":"One-off prompts, no context reuse. Worst-case burn.","cache_hit_rate":0},"mixed":{"label":"Mixed (40% cache)","description":"Interactive coding with some context reuse. Typical default.","cache_hit_rate":0.4},"warm":{"label":"Warm (70% cache)","description":"Sustained agentic session. AA's industry-standard 7:2:1 blend assumes this rate. danielvaughan.com's Codex CLI model uses 70%.","cache_hit_rate":0.7},"hot":{"label":"Hot (90% cache)","description":"Long warmer-pattern sessions with stable system prompts and reused context. vsits.co documented 99% achievable on Claude Code Max subscription with optimization.","cache_hit_rate":0.9}},"default_workload":"mixed","cache_methodology":"Cache discount = (cache_hit_price / input_price). Lower = better discount. Effective burn = (1 - cache_hit_rate) × raw_burn + cache_hit_rate × raw_burn × cache_discount. Per-provider cache discounts: Anthropic 10% (90% off), OpenAI 25% on flagship/10% on 5.4 family, DeepSeek V4-Pro 0.83% (99.2% off — most aggressive in industry), DeepSeek V4-Flash 2%, MiniMax 52%, Mistral &amp; Grok no published cache pricing (modeled as 100% = no discount), Google ~25% plus per-hour storage fee (not modeled). IMPORTANT CAVEAT on Anthropic subscriptions: docs say cache reads count at 10% rate, but GitHub issue anthropics/claude-code#24147 reports cache reads burning quota at full rate on Max subscriptions. Treat the displayed cache benefit as accurate for API usage; subscription quota behavior is contested. Cache writes (1.25× / 2× of input rate on Anthropic) are not modeled — amortize across many reads in steady-state."},"sources":[{"n":1,"category":"Anthropic","title":"Anthropic news","url":"https://www.anthropic.com/news"},{"n":2,"category":"Anthropic","title":"Anthropic Claude model overview","url":"https://docs.claude.com/en/docs/about-claude/models/overview"},{"n":3,"category":"Anthropic","title":"Anthropic pricing","url":"https://www.anthropic.com/pricing"},{"n":4,"category":"OpenAI","title":"OpenAI news","url":"https://openai.com/news/"},{"n":5,"category":"OpenAI","title":"OpenAI model catalog","url":"https://platform.openai.com/docs/models"},{"n":6,"category":"OpenAI","title":"OpenAI API pricing","url":"https://openai.com/api/pricing/"},{"n":7,"category":"Google","title":"Google DeepMind blog","url":"https://blog.google/technology/google-deepmind/"},{"n":8,"category":"Google","title":"Gemini API models","url":"https://ai.google.dev/gemini-api/docs/models"},{"n":9,"category":"Google","title":"Gemini API pricing","url":"https://ai.google.dev/pricing"},{"n":10,"category":"Mistral","title":"Mistral news","url":"https://mistral.ai/news/"},{"n":11,"category":"Mistral","title":"Mistral La Plateforme pricing","url":"https://mistral.ai/products/la-plateforme#pricing"},{"n":12,"category":"xAI","title":"xAI blog (Grok)","url":"https://x.ai/blog"},{"n":13,"category":"DeepSeek","title":"DeepSeek API docs","url":"https://api-docs.deepseek.com/"},{"n":14,"category":"MiniMax","title":"MiniMax news","url":"https://www.minimaxi.com/en/news"},{"n":15,"category":"Benchmarks","title":"SWE-bench (Verified)","url":"https://www.swebench.com/","note":"SWE-bench Verified is widely considered contaminated; treat with skepticism."},{"n":16,"category":"Benchmarks","title":"SWE-bench Pro leaderboard","url":"https://scale.com/leaderboard/swe_bench_pro","note":"Preferred trustworthy code-agent benchmark."},{"n":17,"category":"Benchmarks","title":"LiveCodeBench","url":"https://livecodebench.github.io/"},{"n":18,"category":"Benchmarks","title":"Artificial Analysis (cross-provider pricing + benchmarks)","url":"https://artificialanalysis.ai/"},{"n":19,"category":"Benchmarks","title":"LMArena (chatbot arena)","url":"https://lmarena.ai/"},{"n":20,"category":"Open-weight","title":"Hugging Face — trending models","url":"https://huggingface.co/models?sort=trending"},{"n":21,"category":"Open-weight","title":"Hugging Face blog","url":"https://huggingface.co/blog"},{"n":22,"category":"Open-weight","title":"r/LocalLLaMA (community signal)","url":"https://www.reddit.com/r/LocalLLaMA/"},{"n":23,"category":"Harnesses","title":"Claude Code (anthropics/claude-code)","url":"https://github.com/anthropics/claude-code"},{"n":24,"category":"Harnesses","title":"OpenCode (sst/opencode)","url":"https://github.com/sst/opencode"},{"n":25,"category":"Harnesses","title":"Codex CLI (openai/codex)","url":"https://github.com/openai/codex"},{"n":26,"category":"Harnesses","title":"Gemini CLI (google-gemini/gemini-cli)","url":"https://github.com/google-gemini/gemini-cli"},{"n":27,"category":"Harnesses","title":"Aider (Aider-AI/aider)","url":"https://github.com/Aider-AI/aider"},{"n":28,"category":"Harnesses","title":"Cline (cline/cline)","url":"https://github.com/cline/cline"},{"n":29,"category":"Harnesses","title":"Roo Code (RooVetGit/Roo-Code)","url":"https://github.com/RooVetGit/Roo-Code"},{"n":30,"category":"Harnesses","title":"Goose (block/goose)","url":"https://github.com/block/goose"},{"n":31,"category":"Harnesses","title":"Cursor changelog","url":"https://cursor.com/changelog"},{"n":32,"category":"Harnesses","title":"Windsurf changelog","url":"https://windsurf.com/changelog"},{"n":33,"category":"Hardware","title":"Vast.ai (spot GPU pricing — A6000, H100)","url":"https://vast.ai/"},{"n":34,"category":"Hardware","title":"Apple MacBook Pro (M-series)","url":"https://www.apple.com/shop/buy-mac/macbook-pro"},{"n":35,"category":"Policy","title":"Anthropic OAuth third-party restrictions (Apr 4, 2026)","url":"https://www.anthropic.com/news"},{"n":36,"category":"Policy","title":"OpenAI Codex token-based pricing migration (Apr 2, 2026)","url":"https://openai.com/news/"},{"n":37,"category":"Aggregators","title":"Hacker News (AI tags)","url":"https://news.ycombinator.com/"},{"n":38,"category":"Benchmarks","title":"AIME (math competition benchmark)","url":"https://aimeproblems.com/","note":"Used to anchor Reasoning axis ratings."},{"n":39,"category":"Benchmarks","title":"MMLU / MMLU-Pro (knowledge benchmark)","url":"https://github.com/hendrycks/test","note":"Used to anchor Knowledge axis ratings."},{"n":40,"category":"Benchmarks","title":"tau-bench (tool-use + agentic behaviour)","url":"https://github.com/sierra-research/tau-bench","note":"Used to anchor Agentic axis ratings."},{"n":41,"category":"Nvidia","title":"build.nvidia.com","url":"https://build.nvidia.com/"},{"n":42,"category":"Nvidia","title":"Hugging Face — Nvidia","url":"https://huggingface.co/nvidia"},{"n":43,"category":"Nvidia","title":"Nvidia blogs","url":"https://blogs.nvidia.com/"},{"n":44,"category":"Benchmarks","title":"SWE-bench Pro public leaderboard (Scale)","url":"https://scale.com/leaderboard/swe_bench_pro_public"},{"n":45,"category":"Benchmarks","title":"Scale Labs","url":"https://labs.scale.com/"},{"n":46,"category":"Benchmarks","title":"SWE-Rebench","url":"https://swe-rebench.com/"},{"n":47,"category":"Benchmarks","title":"Terminal-Bench 2.0","url":"https://tbench.ai/leaderboard"},{"n":48,"category":"Aggregators","title":"OpenRouter","url":"https://openrouter.ai/"},{"n":49,"category":"Aggregators","title":"LLM-Stats","url":"https://llm-stats.com/"},{"n":50,"category":"Benchmarks","title":"Artificial Analysis Intelligence Index","url":"https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index"},{"n":51,"category":"Harnesses","title":"Claude Code changelog","url":"https://code.claude.com/docs/en/changelog"},{"n":52,"category":"Harnesses","title":"Cursor pricing","url":"https://cursor.com/pricing"},{"n":53,"category":"Harnesses","title":"Windsurf changelog (Devin Desktop)","url":"https://windsurf.com/changelog"},{"n":54,"category":"Harnesses","title":"Gemini CLI changelogs","url":"https://geminicli.com/docs/changelogs"},{"n":55,"category":"Harnesses","title":"Codex CLI releases","url":"https://github.com/openai/codex/releases"},{"n":56,"category":"Google","title":"DeepMind models","url":"https://deepmind.google/models"},{"n":57,"category":"Mistral","title":"Mistral pricing","url":"https://mistral.ai/pricing"},{"n":58,"category":"Google","title":"Gemini API pricing (canonical)","url":"https://ai.google.dev/gemini-api/docs/pricing"},{"n":59,"category":"MiniMax","title":"Artificial Analysis — MiniMax M2.7","url":"https://artificialanalysis.ai/models/minimax-m2-7"},{"n":60,"category":"Moonshot","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3"},{"n":61,"category":"Moonshot","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models"},{"n":62,"category":"OpenAI","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/"},{"n":63,"category":"Hosting","title":"Runpod GPU cloud pricing","url":"https://www.runpod.io/pricing"},{"n":64,"category":"Hosting","title":"Runpod July 2024 GPU price changes","url":"https://www.runpod.io/blog/runpod-slashes-gpu-prices-more-power-less-cost-for-ai-builders"},{"n":65,"category":"Hosting","title":"Contabo GPU Cloud configurations and pricing","url":"https://contabo.com/en/gpu-cloud/"},{"n":66,"category":"Hosting","title":"Infomaniak Public Cloud pricing","url":"https://www.infomaniak.com/en/hosting/public-cloud/prices"},{"n":67,"category":"Hosting","title":"Infomaniak GPU flavor documentation","url":"https://docs.infomaniak.cloud/compute/instances/flavors/"},{"n":68,"category":"Hosting","title":"Hyperstack GPU cloud pricing","url":"https://www.hyperstack.cloud/"},{"n":69,"category":"Hosting","title":"Vast.ai marketplace pricing methodology","url":"https://docs.vast.ai/guides/instances/pricing"},{"n":70,"category":"Harnesses","title":"Codex CLI 0.144.6 release","url":"https://github.com/openai/codex/releases/tag/rust-v0.144.6"},{"n":71,"category":"Harnesses","title":"Gemini CLI 0.51.0 release","url":"https://github.com/google-gemini/gemini-cli/releases/tag/v0.51.0"},{"n":72,"category":"Harnesses","title":"OpenCode 1.18.4 release","url":"https://github.com/anomalyco/opencode/releases/tag/v1.18.4"},{"n":73,"category":"Harnesses","title":"Cline 4.0.10 release","url":"https://github.com/cline/cline/releases/tag/v4.0.10"},{"n":74,"category":"Harnesses","title":"Goose 1.43.0 release","url":"https://github.com/aaif-goose/goose/releases/tag/v1.43.0"},{"n":75,"category":"Moonshot","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing"},{"n":76,"category":"Swiss AI Initiative","title":"Apertus v1.1 4B Instruct model card","url":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"n":77,"category":"Google","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"n":78,"category":"Google","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"n":79,"category":"Google","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"n":80,"category":"Google","title":"Gemini API release notes","url":"https://ai.google.dev/gemini-api/docs/changelog"},{"n":81,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash"},{"n":82,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"},{"n":83,"category":"OpenAI","title":"GPT-5.6 launch and benchmark table","url":"https://openai.com/index/gpt-5-6/"},{"n":84,"category":"Google","title":"Gemini API Managed Agents update","url":"https://blog.google/innovation-and-ai/technology/developers-tools/expanding-managed-agents-gemini-api/"},{"n":85,"category":"Benchmarks","title":"Scale Labs SWE-bench Pro public leaderboard","url":"https://labs.scale.com/api/pdf/leaderboard/swe_bench_pro_public","note":"Public leaderboard values use their own model versions and evaluation setup; do not merge them mechanically with vendor launch tables."},{"n":86,"category":"Anthropic","title":"What's new in Claude Opus 5","url":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5","note":"Official launch specifications, pricing, availability, and migration behaviour."},{"n":87,"category":"Benchmarks","title":"Artificial Analysis — Claude Opus 5","url":"https://artificialanalysis.ai/models/claude-opus-5","note":"Independent effort-specific intelligence, latency, throughput, and price analysis."},{"n":88,"category":"Thinking Machines Lab","title":"Inkling model card","url":"https://thinkingmachines.ai/model-card/inkling/"},{"n":89,"category":"OpenAI","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra"},{"n":90,"category":"OpenAI","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna"},{"n":91,"category":"DeepSeek","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","note":"Vendor-reported benchmark results use DeepSeek Harness minimal mode at max effort; DSBench results are internal."},{"n":92,"category":"DeepSeek","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/"},{"n":93,"category":"OpenAI","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","note":"Official July 30 pricing, paid-subscription credit, and Sol API Fast mode details."},{"n":94,"category":"OpenAI","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","note":"Official August 6 ChatGPT product update; it does not change the API model contract."}],"provider_sources":{"Anthropic":[1,2,3,86,87],"OpenAI":[4,5,6,62,83,89,90,93,94],"Google":[7,8,9,58,77,78,79,80,81,82,84],"Mistral":[10,11],"xAI":[12],"DeepSeek":[13,91,92],"MiniMax":[14],"Meta":[20,21],"Alibaba":[20,21],"Moonshot":[20,21,60,61,75],"Zhipu":[20,21],"Nvidia":[41,42,43],"Cohere":[20,21],"Swiss AI Initiative":[76],"SubQ":[],"Thinking Machines Lab":[88]},"section_sources":{"models":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,60,61,62,76,77,78,79,80,81,82,86,87,89,90,91,92],"harnesses":[23,24,25,26,27,28,29,30,31,32,35,51,52,53,54,55,70,71,72,73,74],"self_hosting":[20,21,22,33,34,63,64,65,66,67,68,69,76],"strategy":[1,3,4,6,16,18,33,35,36]},"capabilities":{"schema_version":"1.0","axes":[{"key":"coding","label":"Coding","short":"Code","effort_sensitivity":0.5},{"key":"reasoning","label":"Reasoning &amp; Architecture","short":"R&amp;A","effort_sensitivity":0.6},{"key":"knowledge","label":"Knowledge &amp; Research","short":"K&amp;R","effort_sensitivity":0.05},{"key":"comms","label":"Communication &amp; Docs","short":"Comms","effort_sensitivity":0.2},{"key":"multimodal","label":"Multimodal","short":"MM","effort_sensitivity":0.05},{"key":"agentic","label":"Agentic","short":"Agent","effort_sensitivity":0.4}],"level_labels":{"coding":{"1":"Snippet","2":"Standard","3":"Cross-file","4":"Hard","5":"Frontier"},"reasoning":{"1":"Apply known","2":"Multi-step","3":"Cross-cutting","4":"Novel system","5":"Research-grade"},"knowledge":{"1":"Recall","2":"Contextual","3":"Cross-domain","4":"Frontier","5":"Original"},"comms":{"1":"Grammatical","2":"Structured","3":"Tutorial","4":"Editorial","5":"Publishable"},"multimodal":{"1":"Text only","2":"Image-in","3":"Image reasoning","4":"Video/audio","5":"Cross-modal gen"},"agentic":{"1":"Single-turn","2":"Tool calls","3":"Plan coherence","4":"Self-correcting","5":"Long-horizon"}},"focus_presets":{"balanced":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"coding-focused":{"coding":2,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"architecture-focused":{"coding":1,"reasoning":2,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"research-focused":{"coding":1,"reasoning":1,"knowledge":2,"comms":1,"multimodal":1,"agentic":1},"writing-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":2,"multimodal":1,"agentic":1},"multimodal-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":2,"agentic":1},"agentic-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":2}},"default_focus":"balanced","methodology":"1-5 levels per axis; Coding rated against SWE-Pro / SWE-Verified evidence; Reasoning against AIME / LiveCodeBench / public hard-reasoning evals; Multimodal against published modality support (image/audio/video in/out); Agentic rated holding harness constant at Claude-Code-class baseline (plan coherence, tool-call quality, self-correction, calibrated stopping, refusal hygiene); Knowledge from model-card claims + community-validated state-of-art awareness; Comms from prose-evaluation rounds + structural quality. Ratings carry citation; see src/dashboard-context.md for the full rubric.","effort_formula":"effective(axis, effort) = max(0, ceiling[axis] × (1 − effort_sensitivity[axis] × (1 − radar_effort_factor[effort]))). radar_effort_factor[medium] = 1.00 → primary polygon equals reference polygon shape at medium effort for the same model. Lower efforts shrink the polygon by sensitivity-weighted amounts; higher efforts grow it and may push vertices OUTSIDE the outer level-50 ring — the ring is a rubric marker, not a hard cap. Vertices that exceed 50 are drawn in accent-hot to signal 'boosted above medium-effort baseline'.","radar_effort_factors":{"minimal":0.5,"low":0.75,"medium":1,"high":1.08,"xhigh":1.15,"max":1.22},"scale":{"rubric_min":1,"rubric_max":50,"visual_max":60,"band_size":10,"boost_band":"51–60","interpretation":"capability_levels[axis] = the model's effective capability at MEDIUM effort. Briefings rate from 1 to 50. The radar's visual scale extends to 60 to accommodate the 'effort-boost band' (51–60) — values that fall there are formula-derived from extra reasoning effort, never directly rated. The outer rubric ring at 50 is visually emphasized as 'current frontier'.","bands":{"1-10":"Snippet","11-20":"Standard","21-30":"Cross-file","31-40":"Hard","41-50":"Frontier","51-60":"Effort-boost band (derived, not rated)"}}},"report_metrics":{"schema_version":"1.0","researched_at":"2026-08-01","currency":"USD","reference_options":["fable-5","gpt-5.6-sol","gpt-5.6-terra","gpt-5.6-luna","opus-4.8","gpt-5.5","gpt-5.5-pro","gpt-5.5-instant","kimi-k3","gemini-3.6-flash","gemini-3.5-flash-lite"],"default_reference":"fable-5","methodology":{"benchmark_scale":"Each benchmark is divided by benchmark.max_value and expressed on a 0-100 scale. The quality composite is the weighted arithmetic mean of available normalized benchmarks; it is shown only when coverage is at least 0.50.","quality_weights":{"aaii_v4_1":0.2,"coding_agent_index_v1_1":0.2,"swe_bench_pro":0.2,"deep_swe_v1_1":0.15,"terminal_bench_2_1":0.15,"agents_last_exam":0.1},"speed_score":"100 * sqrt(task_speed_index / max_task_speed_index). task_speed_index is 100 * reference_time / model_time. It is not output tokens per second and is comparable only within the cited evaluation setup.","cost_score":"10 + 90 * ln(max_blended_price / blended_price) / ln(max_blended_price / min_blended_price). Blended price is 0.30 * input_price + 0.70 * output_price per million tokens. Higher cost score is better.","scq_compound":"Geometric mean of quality_score, speed_score and cost_score. Geometric mean prevents one category from fully compensating for a weak category. Null when quality coverage is below 0.50.","capability_compound":"Weighted arithmetic mean of the six capability axes. Balanced uses weight 1 for each axis; a selected focus uses weight 2.5 for that axis and 1 for every other axis.","reference_quality":"Each model stores or derives a Fable-anchored quality value. Displayed quality = 100 * model_anchor / selected_reference_anchor, so the selected reference is always exactly 100. New-suite composites use only explicitly overlapping evaluations; legacy rows fall back to SWE-Bench Pro. Unknown evidence remains null, never zero.","burn":"Raw blended-price units are multiplied by an effort factor and cache factor, then divided by the selected reference model at medium effort. Cache factor = (1-hit_rate) + hit_rate*cache_read_ratio.","missing_data":"Keep unknown values null. Display insufficient comparable evidence rather than zero. Detailed benchmark cells remain version-specific and may be empty even when a reference-relative quality anchor exists from a separate documented comparison set."},"market_signals":[{"id":"frontier_quality","label":"Frontier quality index","explanation":"Best broad intelligence score published on Artificial Analysis Intelligence Index v4.1. Higher is better; the benchmark version must remain fixed across the series.","unit":"index","direction":"up","current":59.9,"history":[{"date":"2026-01-01","value":48.2},{"date":"2026-02-01","value":50.1},{"date":"2026-03-01","value":52.7},{"date":"2026-04-01","value":55.7},{"date":"2026-05-01","value":55.7},{"date":"2026-06-01","value":59.9},{"date":"2026-07-01","value":59.9}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"coding_quality","label":"Coding-agent frontier","explanation":"Best score on Artificial Analysis Coding Agent Index v1.1. Higher means better end-to-end coding-agent performance, not just code completion.","unit":"index","direction":"up","current":80,"history":[{"date":"2026-01-01","value":66.1},{"date":"2026-02-01","value":69.8},{"date":"2026-03-01","value":72.5},{"date":"2026-04-01","value":76.4},{"date":"2026-05-01","value":77.2},{"date":"2026-06-01","value":77.2},{"date":"2026-07-01","value":80}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"task_speed","label":"Coding task speed increase","explanation":"Fastest current end-to-end coding task-rate index, where Fable 5 is fixed at 100. A value of 320 means an estimated 3.2 times as many comparable tasks per unit time; it is not tokens per second.","unit":"Fable=100","direction":"up","current":320,"history":[{"date":"2026-01-01","value":100},{"date":"2026-02-01","value":112},{"date":"2026-03-01","value":126},{"date":"2026-04-01","value":145},{"date":"2026-05-01","value":180},{"date":"2026-06-01","value":180},{"date":"2026-07-01","value":320}],"source":"https://openai.com/index/gpt-5-6/","history_note":"Pre-July points are frozen planning estimates from prior report snapshots and should not be interpreted as one controlled longitudinal benchmark."},{"id":"blended_frontier_price","label":"Lowest frontier blended API price","explanation":"Lowest 30% input / 70% output list-price blend among models meeting the report's frontier quality floor. Lower is better; batch and caching are excluded.","unit":"$/Mtok","direction":"down","current":0.9,"history":[{"date":"2026-01-01","value":18},{"date":"2026-02-01","value":16.5},{"date":"2026-03-01","value":14},{"date":"2026-04-01","value":11.4},{"date":"2026-05-01","value":9},{"date":"2026-06-01","value":9},{"date":"2026-07-01","value":4.5},{"date":"2026-08-01","value":0.9}],"source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"open_weight_quality","label":"Open-weight quality vs selected reference","explanation":"Best open-weight model relative to the selected reference. Stored history is Fable-anchored and rebased in the browser whenever the reference changes. July uses Kimi K3's geometric mean across 14 overlapping launch-suite values; vendor-reported values require independent reproduction.","unit":"%","direction":"up","current":99.2,"history":[{"date":"2026-01-01","value":82},{"date":"2026-02-01","value":85},{"date":"2026-03-01","value":88},{"date":"2026-04-01","value":91},{"date":"2026-05-01","value":94},{"date":"2026-06-01","value":96},{"date":"2026-07-01","value":99.2}],"history_note":"Earlier points are frozen report estimates; July uses Moonshot’s K3 launch suite and is not a single controlled longitudinal benchmark.","source":"https://www.kimi.com/blog/kimi-k3"},{"id":"output_throughput","label":"Documented API output throughput","explanation":"Fastest provider-documented general API output rate in the tracked set. Unit is generated tokens per second; it is distinct from time to first token and end-to-end task speed.","unit":"tok/s","direction":"up","current":490,"history":[{"date":"2026-01-01","value":110},{"date":"2026-02-01","value":130},{"date":"2026-03-01","value":140},{"date":"2026-04-01","value":180},{"date":"2026-05-01","value":200},{"date":"2026-06-01","value":220},{"date":"2026-07-01","value":490}],"history_note":"Provider and independent figures use different serving and context conditions. The latest point is Artificial Analysis's Gemini 3.5 Flash-Lite measurement on Google's first-party API.","source":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"}],"benchmarks":[{"id":"aaii_v4_1","label":"AA Intelligence Index v4.1","max_value":59.9,"unit":"index","good_direction":"high"},{"id":"coding_agent_index_v1_1","label":"AA Coding Agent Index v1.1","max_value":80,"unit":"index","good_direction":"high"},{"id":"swe_bench_pro","label":"SWE-Bench Pro","max_value":80,"unit":"%","good_direction":"high"},{"id":"deep_swe_v1_1","label":"DeepSWE v1.1","max_value":72.7,"unit":"%","good_direction":"high"},{"id":"terminal_bench_2_1","label":"Terminal-Bench 2.1","max_value":91.9,"unit":"%","good_direction":"high","note":"Maximum is GPT-5.6 Sol Ultra; ordinary Sol scores 88.8."},{"id":"agents_last_exam","label":"Agents' Last Exam","max_value":52.7,"unit":"%","good_direction":"high"}],"model_metrics":[{"model_id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","benchmarks":{"aaii_v4_1":59.9,"coding_agent_index_v1_1":77.2,"swe_bench_pro":80,"deep_swe_v1_1":69.7,"terminal_bench_2_1":83.1,"agents_last_exam":40.5},"quality_score":94.9,"quality_coverage":1,"task_speed_index":100,"speed_score":50,"api_input":10,"api_cached_input":1,"api_output":50,"blended_price":38,"cost_score":10,"scq_compound":36.2,"speed_evidence":"Reference index; Anthropic labels comparative latency slower.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","benchmarks":{"aaii_v4_1":58.9,"coding_agent_index_v1_1":80,"swe_bench_pro":64.6,"deep_swe_v1_1":72.7,"terminal_bench_2_1":88.8,"agents_last_exam":52.7},"quality_score":95.4,"quality_coverage":1,"task_speed_index":256,"speed_score":80,"api_input":5,"api_cached_input":0.5,"api_output":30,"blended_price":22.5,"cost_score":22.6,"scq_compound":55.7,"speed_evidence":"OpenAI introduced API Fast mode on July 30, claiming up to 2.5× Standard speed at 2× price with no intelligence change; the base task-speed index remains a separate vendor comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","benchmarks":{"aaii_v4_1":55,"coding_agent_index_v1_1":77.4,"swe_bench_pro":63.4,"deep_swe_v1_1":69.6,"terminal_bench_2_1":87.4,"agents_last_exam":50.4},"quality_score":91.8,"quality_coverage":1,"task_speed_index":300,"speed_score":86.6,"api_input":2,"api_cached_input":0.2,"api_output":12,"blended_price":9,"cost_score":44.6,"scq_compound":70.8,"speed_evidence":"OpenAI reports roughly one-third of Fable 5 task time for the family coding comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","benchmarks":{"aaii_v4_1":51.2,"coding_agent_index_v1_1":74.6,"swe_bench_pro":62.7,"deep_swe_v1_1":67.2,"terminal_bench_2_1":84.7,"agents_last_exam":50.3},"quality_score":88.7,"quality_coverage":1,"task_speed_index":320,"speed_score":89.4,"api_input":0.2,"api_cached_input":0.02,"api_output":1.2,"blended_price":0.9,"cost_score":100,"scq_compound":92.6,"speed_evidence":"OpenAI calls Luna the fastest tier; 320 is a conservative planning index above the cited 300 family comparison and must not be treated as measured tok/s.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","benchmarks":{"deep_swe_v1_1":67.5,"terminal_bench_2_1":88.3},"quality_score":null,"quality_coverage":0.33,"quality_vs_fable":99.2,"task_speed_index":null,"speed_score":null,"api_input":3,"api_cached_input":0.3,"api_output":15,"blended_price":11.4,"cost_score":38.9,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time or output-throughput figure published for K3 at launch.","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"model_id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","benchmarks":{"aaii_v4_1":55.7,"coding_agent_index_v1_1":72.5,"swe_bench_pro":69.2,"deep_swe_v1_1":59,"terminal_bench_2_1":78.9,"agents_last_exam":45.2},"quality_score":87.7,"quality_coverage":1,"task_speed_index":180,"speed_score":67.1,"api_input":5,"api_cached_input":0.5,"api_output":25,"blended_price":19,"cost_score":26.7,"scq_compound":54,"speed_evidence":"Planning index based on moderate vendor latency and optional fast mode; validate in the local harness.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gemini-3.1-pro","name":"Gemini 3.1 Pro Preview","provider":"Google","benchmarks":{"aaii_v4_1":46.5,"coding_agent_index_v1_1":42.7,"swe_bench_pro":54.2,"deep_swe_v1_1":11.8,"terminal_bench_2_1":70.7,"agents_last_exam":32.1},"quality_score":59.8,"quality_coverage":1,"task_speed_index":null,"speed_score":null,"api_input":null,"api_cached_input":null,"api_output":null,"blended_price":null,"cost_score":null,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time measurement in the selected source set.","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro"},{"model_id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","benchmarks":{"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78},"quality_score":77.4,"quality_coverage":0.7,"quality_vs_fable":81.6,"task_speed_index":null,"speed_score":null,"api_input":1.5,"api_cached_input":0.15,"api_output":7.5,"blended_price":5.7,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 304 output tok/s at high thinking; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"model_id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","benchmarks":{"aaii_v4_1":36,"swe_bench_pro":54.2,"terminal_bench_2_1":54},"quality_score":62.5,"quality_coverage":0.55,"quality_vs_fable":65.9,"task_speed_index":null,"speed_score":null,"api_input":0.3,"api_cached_input":0.03,"api_output":2.5,"blended_price":1.84,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 490 output tok/s; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"}],"visualizations":{"bubble":{"x":"cost_score","y":"speed_score","size":"quality_vs_selected_reference","color":"provider","include_when":"quality vs reference, speed score and cost score are all non-null"},"heatmap":{"rows":"model","columns":["quality_vs_selected_reference","speed_score","cost_score"],"color_scale":"sequential_good","domain":[0,100]},"bcg":{"x":"capability_compound","y":"mean(quality_vs_selected_reference, speed_score, cost_score)","color":"provider","quadrants":"medians of currently visible models"}},"capability":{"axes":["coding","reasoning_architecture","knowledge_research","communication_docs","multimodal","agentic"],"focus_multiplier":2.5,"models":[{"model_id":"fable-5","scores":[98,99,96,96,90,99],"capability_compound":96.3,"scq_compound":36.2},{"model_id":"gpt-5.6-sol","scores":[98,97,95,95,96,98],"capability_compound":96.5,"scq_compound":55.7},{"model_id":"gpt-5.6-terra","scores":[95,92,91,92,90,95],"capability_compound":92.5,"scq_compound":70.8},{"model_id":"gpt-5.6-luna","scores":[92,88,86,89,86,91],"capability_compound":88.7,"scq_compound":92.6},{"model_id":"opus-4.8","scores":[92,94,94,95,82,94],"capability_compound":91.8,"scq_compound":54},{"model_id":"gemini-3.1-pro","scores":[78,84,94,84,98,72],"capability_compound":85,"scq_compound":null}],"provenance_note":"Capability scores are editorial rubric ratings synthesized from benchmark and feature evidence, not vendor benchmark results. Keep them separate from benchmark values."},"economics":{"effort_factors":{"none":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4},"cache_hit_presets":{"cold":0,"mixed":0.4,"warm":0.7,"hot":0.9},"models":[{"model_id":"fable-5","base_blended_price":38,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"],"note":"Always-on adaptive thinking; effort is a behavioral control, not a published token multiplier."},{"model_id":"gpt-5.6-sol","base_blended_price":22.5,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-terra","base_blended_price":9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-luna","base_blended_price":0.9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"opus-4.8","base_blended_price":19,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"]},{"model_id":"kimi-k3","base_blended_price":11.4,"cache_read_ratio":0.1,"efforts":["max"],"note":"Always-on reasoning at launch; the max-only effort control is behavioral and not a published token multiplier."}],"efficiency_filters":{"quality_floor_default":75,"provider_default":"all","effort_default":"medium","workload_default":"mixed","sort_default":"scq_compound_desc"}},"hardware_options":[{"id":"local-64gb","name":"64 GB unified-memory workstation","memory_gb":64,"type":"local","best_for":"3B-35B dense or small MoE models at Q4-Q8","cost_note":"Capex varies; compare measured memory bandwidth, not product year."},{"id":"local-128gb","name":"128 GB unified-memory workstation","memory_gb":128,"type":"local","best_for":"Up to roughly 100 GB quantized weights with headroom for KV cache","cost_note":"Portable and quiet; slower than datacenter GPUs for sustained batches."},{"id":"cloud-2xa6000","name":"2 x RTX A6000 48 GB","memory_gb":96,"type":"cloud","best_for":"70B-class dense and mid-size MoE Q4 deployments","cost_note":"Spot price varies by host; record price and interconnect at test time."},{"id":"cloud-h200","name":"1 x H200 141 GB","memory_gb":141,"type":"cloud","best_for":"High-throughput 70B inference and larger quantized MoE models","cost_note":"Use provider quote; hourly rates change frequently."},{"id":"cloud-b200","name":"1 x B200 192 GB","memory_gb":192,"type":"cloud","best_for":"Large-model throughput where software stack supports Blackwell","cost_note":"Availability and hourly rates vary; validate framework support."}],"hardware_model_fit":[{"model_id":"qwen3.6-35b-a3b","hardware_id":"local-64gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":55,"evidence":"planning_estimate"},{"model_id":"qwen3.6-35b-a3b","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"BF16","estimated_output_tps":140,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"local-64gb","fit_status":"unsupported","reason":"Quantized weights exceed usable memory."},{"model_id":"kimi-k2.6-1t","hardware_id":"local-128gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":16,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"cloud-h200","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":36,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-flash-open","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":45,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-h200","fit_status":"unsupported","reason":"Single-device memory is insufficient at the tracked quantization."},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-b200","fit_status":"supported","quant":"Q2_K","estimated_output_tps":14,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"local-64gb","fit_status":"supported","quant":"BF16","estimated_output_tps":38,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"cloud-h200","fit_status":"supported","quant":"BF16","estimated_output_tps":80,"evidence":"planning_estimate"}],"decision_support":{"recommended_stack":[{"route":"hardest_long_horizon","primary":"fable-5","effort":"high","fallback":"gpt-5.6-sol","why":"Use Fable where long-horizon reliability warrants cost and retention policy is acceptable."},{"route":"default_engineering","primary":"gpt-5.6-terra","effort":"medium","fallback":"gpt-5.6-sol","why":"Best balanced default in the current benchmark/cost set."},{"route":"high_volume_subagents","primary":"gpt-5.6-luna","effort":"low","fallback":"gemini-3.5-flash-lite","why":"Luna retains stronger coding-agent evidence; Flash-Lite is the throughput-first fallback for extraction and delegated subtasks."},{"route":"privacy_local","primary":"qwen3.6-35b-a3b","effort":null,"fallback":"deepseek-v4-flash-open","why":"Local-first route where external retention is unacceptable."},{"route":"vision_long_context","primary":"gemini-3.1-pro","effort":"high","fallback":"gpt-5.6-sol","why":"Prefer for multimodal context; validate preview stability before production."}],"routing_rules":["Route by task risk and evidence, not provider family.","Start bulk work on Luna or Terra and escalate only after an explicit verification failure.","Do not send ZDR-required data to Fable 5 because the model requires 30-day retention.","Record model, effort, cache state, region and harness version for every internal speed comparison.","Use self-hosted routes only when the selected hardware row is supported; never infer performance from an empty cell."]},"actions":[{"priority":"P0","owner":"Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","status":"open","order":1},{"priority":"P0","owner":"Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","status":"open","order":2},{"priority":"P0","owner":"Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","status":"open","order":3},{"priority":"P1","owner":"Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","status":"open","order":4},{"priority":"P1","owner":"Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","status":"open","order":5},{"priority":"P2","owner":"Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics.","status":"open","order":6}],"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the price-change history using OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. Added the paid Codex and ChatGPT Work credit reduction, unchanged subscription prices and quota budgets, and Sol API Fast mode (up to 2.5× Standard speed at 2× price)."},{"date":"2026-08-01","tag":"pricing","text":"Refreshed OpenAI's live GPT-5.6 API rate card: Terra is now $2/$12 per million input/output tokens and Luna is $0.20/$1.20, with cached input at $0.20 and $0.02 respectively. Recomputed the 30/70 workload blend, cost-efficiency scores, SCQ compounds and burn baselines. Superseded on August 7 with the official July 30 effective date and full change details."},{"date":"2026-08-01","tag":"benchmark","text":"Recorded DeepSeek V4-Flash-0731's July 31 public-beta evidence: Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in DeepSeek Harness minimal mode. Additional vendor-reported NL2Repo, Cybergym, Toolathlon and Automation Bench results remain narrative evidence rather than new shared columns."},{"date":"2026-07-22","tag":"method","text":"Marked the open-weight quality history as a Fable-anchored source series that is rebased in the browser against the selected reference; fixed task-speed remains explicitly Fable-indexed."},{"date":"2026-07-22","tag":"model","text":"Added Gemini 3.6 Flash and Gemini 3.5 Flash-Lite benchmark, pricing, context and throughput evidence from Google and Artificial Analysis; retained missing task-time fields as null."},{"date":"2026-07-21","tag":"model","text":"Added official Kimi K3 API pricing and derived the documented 30/70 workload blend and cost-efficiency score; no task-speed compound is shown because comparable speed evidence remains unavailable."},{"date":"2026-07-19","text":"Added a dated GPU-hosting price tracker with billing-basis normalization, historical Runpod price changes, current Contabo configurations, and quote-aware Infomaniak options. Replaced stale Vast.ai constants with live-marketplace status."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, context, benchmark records and reference eligibility."},{"date":"2026-07-18","tag":"method","text":"Introduced documented quality, speed, cost, SCQ and six-axis capability composites with explicit missing-data rules."},{"date":"2026-07-18","tag":"hardware","text":"Replaced blank hardware fit speeds with supported/unsupported states and evidence-labelled planning estimates."},{"date":"2026-07-18","tag":"routing","text":"Updated recommended routing to Fable for hardest retained-data work, Terra for default engineering and Luna for high-volume subagents."}],"sources":[{"id":"openai-gpt-5-6","title":"GPT-5.6 launch and evaluations","url":"https://openai.com/index/gpt-5-6/","accessed":"2026-07-18"},{"id":"moonshot-kimi-k3-pricing","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing","accessed":"2026-07-21"},{"id":"openai-models","title":"OpenAI model catalog","url":"https://developers.openai.com/api/docs/models","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-terra","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-luna","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-price-performance","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","accessed":"2026-08-07"},{"id":"openai-gpt-5-6-chat-access","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","accessed":"2026-08-07"},{"id":"deepseek-v4-flash-0731","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","accessed":"2026-08-01"},{"id":"deepseek-v4-pricing","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/","accessed":"2026-08-01"},{"id":"anthropic-models","title":"Claude models overview","url":"https://platform.claude.com/docs/en/about-claude/models/overview","accessed":"2026-07-18"},{"id":"anthropic-retention","title":"Claude API and data retention","url":"https://platform.claude.com/docs/en/manage-claude/api-and-data-retention","accessed":"2026-07-18"},{"id":"google-gemini-3-5","title":"Gemini 3.5 Flash","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash","accessed":"2026-07-18"},{"id":"google-gemini-3-6","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash","accessed":"2026-07-22"},{"id":"google-gemini-3-5-flash-lite","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite","accessed":"2026-07-22"},{"id":"google-flash-july-2026","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/","accessed":"2026-07-22"},{"id":"aa-gemini-3-6-flash","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash","accessed":"2026-07-22"},{"id":"aa-gemini-3-5-flash-lite","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","accessed":"2026-07-22"},{"id":"moonshot-kimi-k3","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3","accessed":"2026-07-18"},{"id":"moonshot-models","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models","accessed":"2026-07-18"},{"id":"openai-gpt-5-5","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/","accessed":"2026-07-18"}],"hosting_prices":{"as_of":"2026-07-19","hours_per_month":730,"methodology":"Preserve provider currency and billing basis. normalized_hourly is a comparison-only derivation for fixed monthly plans; monthly_equivalent multiplies hourly rates by 730. It excludes tax, storage, egress, public IPs, support, discounts, utilization, and model throughput. Append observations instead of replacing them.","trend_policy":"Track the same provider, GPU, service tier, region basis, and currency. A changed SKU starts a new series. Two or more observations produce a trend; one observation is a dated baseline.","offers":[{"id":"runpod-h100-sxm-secure","provider":"Runpod","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$2.99/hr","normalized_hourly":2.99,"monthly_equivalent":2182.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":3.99,"source_n":64},{"date":"2026-07-19","value":2.99,"source_n":63}]},{"id":"runpod-a100-sxm-secure","provider":"Runpod","configuration":"1 x A100 SXM","gpu":"A100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$1.49/hr","normalized_hourly":1.49,"monthly_equivalent":1087.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.94,"source_n":64},{"date":"2026-07-19","value":1.49,"source_n":63}]},{"id":"runpod-l40s-secure","provider":"Runpod","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$0.99/hr","normalized_hourly":0.99,"monthly_equivalent":722.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.19,"source_n":64},{"date":"2026-07-19","value":0.99,"source_n":63}]},{"id":"hyperstack-h100-sxm","provider":"Hyperstack","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$2.40/hr","normalized_hourly":2.4,"monthly_equivalent":1752,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":2.4,"source_n":68}]},{"id":"hyperstack-h200-sxm","provider":"Hyperstack","configuration":"1 x H200 SXM","gpu":"H200 SXM","gpu_count":1,"vram_gb":141,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$3.50/hr","normalized_hourly":3.5,"monthly_equivalent":2555,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":3.5,"source_n":68}]},{"id":"contabo-l40s-monthly","provider":"Contabo","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€751/mo","normalized_hourly":1.0288,"monthly_equivalent":751,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":1.0288,"source_n":65}]},{"id":"contabo-h100-monthly","provider":"Contabo","configuration":"1 x H100","gpu":"H100","gpu_count":1,"vram_gb":80,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€1,838/mo","normalized_hourly":2.5178,"monthly_equivalent":1838,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.5178,"source_n":65}]},{"id":"contabo-h200-nvl-monthly","provider":"Contabo","configuration":"1 x H200 NVL","gpu":"H200 NVL","gpu_count":1,"vram_gb":141,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€2,149/mo","normalized_hourly":2.9438,"monthly_equivalent":2149,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.9438,"source_n":65}]},{"id":"infomaniak-l40s","provider":"Infomaniak","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Switzerland","billing":"Usage-based Public Cloud","price_basis":"Live calculator / availability validation","currency":"CHF","current_price_label":"Calculator / validate","normalized_hourly":null,"monthly_equivalent":null,"price_status":"not_publicly_exposed","checked":"2026-07-19","source_n":66,"history":[]},{"id":"vast-a6000-market","provider":"Vast.ai","configuration":"1 x RTX A6000","gpu":"RTX A6000","gpu_count":1,"vram_gb":48,"region":"Marketplace","billing":"Per second; on-demand, reserved, or interruptible","price_basis":"Live host marketplace","currency":"USD","current_price_label":"Live marketplace","normalized_hourly":null,"monthly_equivalent":null,"price_status":"dynamic_marketplace","checked":"2026-07-19","source_n":69,"history":[]}]}},"model_roster":{"schema_version":"2.0-preview","researched_at":"2026-08-07","default_reference":"fable-5","inclusion_policy":{"default_limit_per_provider":3,"default_scope":"core","summary":"Show a flagship, a balanced or fast model, and one distinctive specialist per provider. Keep siblings and emerging models in expandable extended and watchlist scopes.","promotion_rule":"Promote a watchlist family when at least two signals hold: current official release, meaningful adoption or discussion, differentiated capability or efficiency, and reproducible availability."},"speed_methodology":{"dimensions":["time_to_first_token","output_tokens_per_second","end_to_end_task_time"],"rule":"Compare speed only at the exact model, configuration, provider, region, date, workload, and cache condition. Vendor claims are labelled and unknown values remain unknown.","bands":{"fast":"Designed or measured for low interaction latency","balanced":"General-purpose latency and quality trade-off","deliberate":"Higher reasoning depth or multi-agent execution","unknown":"No comparable evidence"}},"models":[{"id":"fable-5","name":"Claude Fable 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","reference":true,"context":"1M","price":"$10 / $50","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"deliberate","speed_note":"Anthropic comparative latency: slower. No Fable fast mode.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"30-day retention; no zero-data-retention option","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"opus-5","name":"Claude Opus 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Anthropic comparative latency: moderate. Research-preview fast mode claims up to 2.5x higher output throughput at $10 / $50 per MTok.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"Adaptive thinking is on by default; disabling it is supported only through high effort. Independent benchmark results vary materially by effort level and should not be treated as a single generic score.","source":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5"},{"id":"opus-4.8","name":"Claude Opus 4.8","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Optional API fast mode claims up to 2.5x output throughput at premium pricing.","availability":"GA","source":"https://platform.claude.com/docs/en/build-with-claude/fast-mode"},{"id":"sonnet-5","name":"Claude Sonnet 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$3 / $15","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"fast","speed_note":"Anthropic comparative latency: fast.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"haiku-4.5","name":"Claude Haiku 4.5","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"200K","price":"$1 / $5","license":"Proprietary","control":"fixed","levels":[],"default_level":"n/a","speed_band":"fast","speed_note":"Anthropic’s latest verified Haiku and fastest listed Claude; no Haiku 5 is in the current official catalog.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports max reasoning completes comparable AAII work in 61% less time than Fable 5. API Fast mode, introduced July 30, claims up to 2.5x Standard speed at 2x price with unchanged intelligence; no comparable output-tokens/s figure is published.","availability":"GA; ChatGPT, Codex, API","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$2 / $12; cached input $0.20","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 20% from $2.50/$15; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; subscription prices and quota budgets unchanged","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$0.20 / $1.20; cached input $0.02","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"fast","speed_note":"OpenAI positions Luna as the fastest tier and reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 80% from $1/$6; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; rolling out as Free and Go default with unlimited text chats and a Think option","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.5","name":"GPT-5.5","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M API / 400K Codex","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports GPT-5.4-class per-token latency; Codex fast mode is 1.5x throughput at 2.5x cost.","availability":"API, ChatGPT and Codex","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"extended","status":"stable","context":"1M","price":"$30 / $180","license":"Proprietary","control":"reasoning effort","levels":["medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Higher-accuracy tier for difficult work; no comparable public output-throughput figure.","availability":"API and eligible ChatGPT plans","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"harness_bundled","scope":"extended","status":"stable","context":"400K","price":"Included by plan / routed service","license":"Proprietary","control":"fixed","levels":[],"default_level":"provider default","speed_band":"fast","speed_note":"Low-latency ChatGPT route; do not equate service behavior with the reasoning model API.","availability":"ChatGPT","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"preview","context":"1M","price":"$2 / $12","license":"Proprietary","control":"thinking level","levels":["low","medium","high"],"default_level":"high","speed_band":"deliberate","speed_note":"Google warns high thinking may significantly delay the first answer token.","availability":"Preview","source":"https://ai.google.dev/gemini-api/docs/thinking"},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$1.50 / $7.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Artificial Analysis measured about 304 output tok/s at high thinking; Google reports fewer reasoning turns and tool calls than 3.5 Flash.","availability":"GA via Gemini API, AI Studio, Gemini app and Antigravity","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$1.50 / $9","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Still offered; migrate new general Flash workloads to 3.6 Flash for lower output price and stronger agentic performance.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash"},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$0.30 / $2.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Artificial Analysis measured about 490 output tok/s; optimized for high-volume subagents, document parsing and extraction.","availability":"GA via Gemini API, AI Studio and Gemini app rollout","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$0.25 / $1.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Still offered; Gemini 3.5 Flash-Lite is the current high-throughput migration target.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite"},{"id":"grok-4.5","name":"Grok 4.5","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"500K","price":"$2 / $6","license":"Proprietary","control":"reasoning effort","levels":["low","medium","high"],"default_level":"high","speed_band":"fast","speed_note":"xAI reports 80 output tokens/s; vendor measurement conditions are not fully comparable here.","availability":"API and integrations; verify regional availability","source":"https://x.ai/news/grok-4-5"},{"id":"grok-4.20-multi-agent","name":"Grok 4.20 Multi-Agent","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"extended","status":"specialist","context":"1M","price":"$1.25 / $2.50","license":"Proprietary","control":"agent count","levels":["low","medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Levels control a 4- or 16-agent process, not ordinary reasoning depth.","availability":"API","source":"https://docs.x.ai/developers/model-capabilities/text/multi-agent"},{"id":"composer-2.5","name":"Composer 2.5","family":"Composer","provider":"Cursor","region":"US","channel":"harness_bundled","scope":"core","status":"stable","context":"Managed by Cursor","price":"$0.50 / $2.50 standard","license":"Proprietary","control":"speed variant","levels":["standard","fast"],"default_level":"fast","speed_band":"fast","speed_note":"Fast is the default; Cursor documents no public low/medium/high effort selector.","availability":"Cursor only","source":"https://cursor.com/blog/composer-2-5"},{"id":"kimi-k3","name":"Kimi K3","family":"Kimi K3","provider":"Moonshot AI","region":"China","channel":"commercial_api_open_weights_announced","scope":"core","status":"stable","context":"1M","price":"$3 / $15; cached input $0.30","license":"Terms pending weight release","control":"reasoning effort","levels":["max"],"default_level":"max","speed_band":"deliberate","speed_note":"No comparable K3 output-rate figure at launch; max thinking is always enabled.","availability":"API, Kimi, Kimi Work and Kimi Code; weights promised by July 27","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"id":"kimi-k2.7-code","name":"Kimi K2.7 Code","family":"Kimi K2.7","provider":"Moonshot AI","region":"China","channel":"commercial_api","scope":"core","status":"specialist","context":"256K","price":"Current API pricing","license":"Proprietary API","control":"service variant","levels":["standard","high-speed"],"default_level":"standard","speed_band":"fast","speed_note":"High-speed service is documented at about 180 tok/s and up to 260 tok/s for short contexts.","availability":"Kimi API and Kimi Code","source":"https://platform.kimi.ai/docs/models"},{"id":"deepseek-v4","name":"DeepSeek V4","family":"DeepSeek V4","provider":"DeepSeek","region":"China","channel":"open_weight","scope":"core","status":"public_beta","context":"1M","price":"Flash $0.14 / $0.28; cached input $0.0028","license":"MIT","control":"mode","levels":["non-think","think","max"],"default_level":"think","speed_band":"balanced","speed_note":"V4-Flash-0731 reports 82.7 on Terminal-Bench 2.1 at max effort in DeepSeek Harness minimal mode; comparable latency remains unknown.","availability":"V4-Flash-0731 API public beta and weights; future 2x peak-hours pricing announced without an effective date","variants":["Pro","Flash"],"source":"https://api-docs.deepseek.com/updates/"},{"id":"qwen-3.6","name":"Qwen3.6","family":"Qwen3.6","provider":"Alibaba Qwen","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"mode and token budget","levels":["non-thinking","thinking","budget"],"default_level":"thinking","speed_band":"balanced","speed_note":"Budget is provider-native; do not translate it into invented effort labels.","availability":"API and weights","variants":["27B","35B-A3B"],"source":"https://github.com/QwenLM/Qwen3.6"},{"id":"glm-5.2","name":"GLM-5.2","family":"GLM-5","provider":"Z.ai","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"MIT","control":"effort","levels":["adaptive"],"default_level":"adaptive","speed_band":"unknown","speed_note":"Architecture efficiency claims are not a comparable API latency measurement.","availability":"API and weights","source":"https://z.ai/blog/glm-5.2"},{"id":"kimi-k2.5","name":"Kimi K2.5","family":"Kimi K2","provider":"Moonshot AI","region":"China","channel":"open_weight","scope":"extended","status":"deprecated","context":"256K","price":"API and self-hosted","license":"Modified MIT","control":"mode","levels":["instant","thinking"],"default_level":"thinking","speed_band":"balanced","speed_note":"No comparable official throughput measurement; Instant is the lower-latency mode.","availability":"Existing users only; platform sunset August 31, 2026","source":"https://platform.kimi.ai/docs/models"},{"id":"minimax-m3","name":"MiniMax M3","family":"MiniMax M3","provider":"MiniMax","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"Community; non-commercial by default","control":"thinking mode","levels":["disabled","adaptive","enabled"],"default_level":"adaptive","speed_band":"balanced","speed_note":"Vendor claims decode improvements versus M2, not cross-provider latency.","availability":"API and weights","caveat":"License is not permissive open source","source":"https://www.minimax.io/blog/minimax-m3"},{"id":"mistral-small-4","name":"Mistral Small 4","family":"Mistral 3/4","provider":"Mistral AI","region":"EU","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"reasoning effort","levels":["none","high"],"default_level":"none","speed_band":"fast","speed_note":"Vendor claims 40% lower completion time and 3x requests/s versus Small 3.","availability":"API and weights","source":"https://mistral.ai/it/news/mistral-small-4/"},{"id":"llama-4","name":"Llama 4","family":"Llama 4","provider":"Meta","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"10M advertised for Scout","price":"Self-hosted","license":"Llama 4 Community License","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Performance depends on deployment; no comparable official API latency.","availability":"Weights","variants":["Scout","Maverick"],"caveat":"Not OSI-open; EU multimodal and large-platform restrictions apply","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"id":"gemma-3","name":"Gemma 3","family":"Gemma 3","provider":"Google","region":"US","channel":"open_weight","scope":"extended","status":"stable","context":"128K","price":"Self-hosted","license":"Google Gemma Terms","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Deployment dependent; optimized for edge and local use.","availability":"Weights","variants":["1B","4B","12B","27B"],"source":"https://ai.google.dev/gemma/docs/core/model_card_3"},{"id":"gpt-oss","name":"gpt-oss","family":"gpt-oss","provider":"OpenAI","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"128K","price":"Self-hosted","license":"Apache-2.0 plus usage policy","control":"reasoning effort","levels":["low","medium","high"],"default_level":"medium","speed_band":"unknown","speed_note":"Depends on hardware and serving stack.","availability":"Weights","variants":["120b","20b"],"source":"https://openai.com/index/introducing-gpt-oss/"},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","family":"Apertus v1.1","provider":"Swiss AI Initiative","region":"Switzerland","channel":"open_weight","scope":"core","status":"stable","context":"4K","price":"Self-hosted","license":"Apache-2.0","control":"sampling parameters","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"No comparable throughput measurement is published; performance depends on hardware, runtime and quantization.","availability":"Weights; Transformers, vLLM, SGLang and quantized checkpoints","variants":["0.5B","1.5B","4B"],"source":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"id":"inkling","name":"Inkling","family":"Inkling","provider":"Thinking Machines Lab","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"No metered API; open weights + Tinker fine-tuning platform","license":"Apache 2.0","control":"n/a","levels":[],"default_level":null,"speed_band":"unknown","speed_note":"No published throughput figure; not yet independently benchmarked for this roster.","availability":"Open weights (Hugging Face), Tinker fine-tuning platform, third-party inference providers","source":"https://thinkingmachines.ai/model-card/inkling/"}],"watchlist":[{"name":"Step 3.5 Flash","provider":"StepFun","region":"China","license":"Apache-2.0","signal":"Official 100-300 output tok/s claim","source":"https://github.com/stepfun-ai/Step-3.5-Flash"},{"name":"MiMo-V2-Flash","provider":"Xiaomi","region":"China","license":"Apache-2.0","signal":"Thinking toggle and generation-speed claim","source":"https://github.com/XiaomiMiMo/MiMo-V2-Flash"},{"name":"Seed-OSS-36B","provider":"ByteDance","region":"China","license":"Apache-2.0","signal":"512K context and controllable thinking budget","source":"https://github.com/ByteDance-Seed/seed-oss"},{"name":"Hy3 Preview","provider":"Tencent","region":"China","license":"Tencent Hy Community License","signal":"Preview multimodal MoE family","source":"https://github.com/Tencent-Hunyuan/Hy3-preview"}]}}&lt;/script&gt;
&lt;section class="tab-panel" id="tab-harnesses"&gt;
 &lt;h2 class="sect-head"&gt;Agent harness landscape&lt;/h2&gt;
 &lt;p class="sect-sub"&gt;Two workflow modes: supervised (you wait) and autonomous (overnight). Anthropic OAuth blocked third-party tools on April 4, 2026 — affected harnesses marked with *.&lt;/p&gt;</description></item><item><title>Project Status</title><link>https://projectious-work.github.io/ai-market-research/docs/project-status/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/docs/project-status/</guid><description>&lt;p&gt;&lt;strong&gt;Lifecycle: active.&lt;/strong&gt; The market roster and published dashboard are actively
maintained. Historical snapshots, when published, remain available for
provenance, but only the current dashboard and latest tagged release are
supported.&lt;/p&gt;
&lt;p&gt;For the live revision, archived-snapshot policy, and release links, see
&lt;a href="https://projectious-work.github.io/ai-market-research/revisions/"&gt;Revisions&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="what-this-is"&gt;What this is&lt;/h2&gt;
&lt;p&gt;Signal Room is inspectable decision support: a manually rebuilt, single static
dashboard that tracks model rosters, provider-native reasoning configurations,
speed evidence, agent harnesses, and self-hosting economics, with evidence
classes and source links kept visible throughout.&lt;/p&gt;</description></item><item><title>03 Infrastructure</title><link>https://projectious-work.github.io/ai-market-research/report/03-infrastructure/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/report/03-infrastructure/</guid><description>&lt;div class="sr-report-scope"&gt;
&lt;script id="market-data" type="application/json"&gt;{"meta":{"generated_at":"2026-08-07T00:00:00Z","reference_default":"fable-5","report_metrics_file":"data/report-metrics.json"},"executive_summary":{"models":["**Claude Opus 5 is now generally available.** Anthropic's new `claude-opus-5` targets complex agentic coding and enterprise work with a 1M-token context window, 128K maximum output, adaptive thinking by default, and unchanged $5/$25 per MTok base pricing. Independent benchmark values are not yet recorded here.","**Google reset the Flash price-performance curve on July 21.** Gemini 3.6 Flash is GA at $1.50/$7.50 per million tokens with stronger coding and agentic results than 3.5 Flash, while Gemini 3.5 Flash-Lite reaches roughly 490 output tok/s at $0.30/$2.50 for high-volume subagents and extraction.","**Kimi K3 resets the open-weight frontier.** Moonshot reports a 2.8T sparse MoE, native vision, 1M context, and 99% of Fable 5 across 14 overlapping launch-suite evaluations; weights are promised by July 27.","**GPT-5.6 Luna and Terra became materially cheaper on July 30.** Luna fell 80% to $0.20/$1.20 per MTok and Terra 20% to $2/$12; their paid Codex and ChatGPT Work usage also consumes fewer credits, while subscription prices and quota budgets did not change.","**GPT-5.6 now spans Sol, Terra and Luna**, giving leaders a deliberate capability, balanced, and high-throughput ladder under one family. Sol API Fast mode replaces Priority Processing: OpenAI claims up to 2.5× Standard speed at 2× price, with unchanged intelligence.","**Luna is expanding beyond API routing.** OpenAI says it becomes the default for Free and Go users, with unlimited text chats and a higher-reasoning Think option rolling out subject to abuse guardrails. This is a ChatGPT product update, not an API capability change.","**Keep benchmark tables separate by harness.** OpenAI's GPT-5.6 launch results and Scale's public SWE-bench Pro leaderboard use different model versions and evaluation setups; use each as evidence, not as one directly rankable series.","**GPT-5.5 remains active in Codex and the API.** It is retained as a compatibility and portfolio option rather than being hidden by the newer 5.6 family.","**Claude Haiku 4.5 remains Anthropic’s latest verified Haiku.** No official Haiku 5 listing was found in Anthropic’s current model catalog."],"harnesses":["**Separate model capability from harness capability.** Tool execution, context management, isolation, and observability can dominate real workflow outcomes.","**Subscription access is not a production routing contract.** Validate API, credit-pool, and third-party harness policies before standardizing an operating model.","**Maintain at least one portable fallback path** across providers for high-value workflows and operational incidents.","**Gemini Managed Agents now add background execution, remote MCP, custom functions and credential refresh.** That makes Google a more credible managed-agent control plane, but it does not substitute for workload-specific evaluation."],"self_hosting":["**Open weights are now a strategic option, not only a cost play.** Kimi K3, DeepSeek, Qwen, Nemotron, Mistral, and Llama cover different sovereignty and specialization needs.","**Do not compare self-hosting at zero token cost.** Include accelerator rental or depreciation, power, utilization, serving staff, and measured throughput.","**Pilot against a defined workload and hardware envelope** before treating advertised context or parameter scale as deployable capacity."],"strategy":["**Run a portfolio, not a winner-takes-all model standard.** Reserve frontier reasoning for high-value decisions and route routine work to measured fast or efficient tiers.","**Instrument quality, latency, retries, and total workflow cost together.** Token price alone is not an operating metric.","**Review the portfolio quarterly and after major releases**, with explicit retirement, security, and fallback criteria."]},"headline_stats":[{"id":"frontier_count","label":"Models tracked","value":35,"unit":"","delta":"Current, fast and retained compatibility models","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_weight","label":"Open-weight","color":"#16866f"}],"history":[{"date":"2026-05-18","proprietary":20,"open_weight":3},{"date":"2026-07-22","proprietary":26,"open_weight":8},{"date":"2026-07-25","proprietary":27,"open_weight":8}],"note":"Release-tag roster snapshots; announced open-weight releases are counted with open-weight models."}},{"id":"best_open_pct","label":"Best open-weight vs selected reference","value":99,"unit":"%","delta":"Kimi K3; geometric mean across 14 overlapping vendor evals","delta_dir":"up","signal_id":"open_weight_quality"},{"id":"cheapest_frontier_api","label":"Lowest frontier blended API price","value":0.9,"unit":"$/Mtok","delta":"GPT-5.6 Luna; 30% input / 70% output blend","delta_dir":"down","signal_id":"blended_frontier_price"},{"id":"fastest_task_rate","label":"Fastest task-rate index","value":"3.2×","unit":"","delta":"Fable 5 = 1.0×; planning index, not tokens/second","delta_dir":"up","signal_id":"task_speed"},{"id":"max_context","label":"Largest usable context","value":"10M","unit":"tok","delta":"Llama 4 Scout; deployment constraints still apply","delta_dir":"up","trend":[{"date":"2026-01-01","value":1},{"date":"2026-04-01","value":1},{"date":"2026-07-01","value":10}]},{"id":"harness_count","label":"Harnesses tracked","value":16,"unit":"","delta":"Commercial and open agent environments","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_source","label":"Open source","color":"#3d78c5"}],"history":[{"date":"2026-05-18","proprietary":3,"open_source":11},{"date":"2026-07-22","proprietary":6,"open_source":10}],"note":"Release-tag roster snapshots classified from each harness license."}},{"id":"documented_output_speed","label":"Fastest documented API output","value":490,"unit":"tok/s","delta":"Gemini 3.5 Flash-Lite; Artificial Analysis first-party API measurement","delta_dir":"up","signal_id":"output_throughput"},{"id":"quality_coverage","label":"Models with quality evidence","value":"31/35","unit":"","delta":"Unknown remains unknown; no zero-value substitution","delta_dir":"neutral","trend":[{"date":"2026-01-01","value":18},{"date":"2026-04-01","value":24},{"date":"2026-07-01","value":31}]}],"trends":{"best_open_vs_opus":{"label":"Best open-weight % vs Opus 4.7","history":[{"date":"2025-11-01","value":68},{"date":"2025-12-01","value":72},{"date":"2026-01-01","value":78},{"date":"2026-02-01","value":80},{"date":"2026-03-01","value":83},{"date":"2026-04-01","value":86},{"date":"2026-05-01","value":88},{"date":"2026-06-01","value":90}]},"median_frontier_output_price":{"label":"Median frontier $/Mtok output","history":[{"date":"2025-11-01","value":18},{"date":"2025-12-01","value":17},{"date":"2026-01-01","value":15.5},{"date":"2026-02-01","value":14.5},{"date":"2026-03-01","value":13},{"date":"2026-04-01","value":12.5},{"date":"2026-05-01","value":12},{"date":"2026-06-01","value":11}]},"models_released_per_month":{"label":"Notable model releases per month","history":[{"date":"2025-11-01","value":3},{"date":"2025-12-01","value":4},{"date":"2026-01-01","value":5},{"date":"2026-02-01","value":6},{"date":"2026-03-01","value":4},{"date":"2026-04-01","value":7},{"date":"2026-05-01","value":5},{"date":"2026-06-01","value":8}]}},"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the GPT-5.6 pricing history from OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. The report now records lower paid Codex and ChatGPT Work credit consumption, unchanged subscription prices and quota budgets, Sol API Fast mode (up to 2.5× Standard speed at 2× price), and the August 6 ChatGPT Luna access expansion."},{"date":"2026-08-01","tag":"pricing","text":"Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to $0.20/$1.20 from OpenAI's live model pages. Recomputed the tracked 30/70 workload blends and effort-burn ratios. Superseded on August 7 with OpenAI's published July 30 effective date and full price-change details."},{"date":"2026-08-01","tag":"model","text":"Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output, thinking/non-thinking modes, $0.14/$0.28 per-million pricing, $0.0028 cached input, and vendor-reported agent results including Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2."},{"date":"2026-07-28","tag":"model","text":"Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B active) open-weight (Apache 2.0) multimodal MoE with 1M-token context, released 2026-07-15. No first-party per-token API pricing is published (monetized via the Tinker fine-tuning platform); benchmark scores are published on the model card but not yet normalized into this roster's comparable set, so quality/speed/cost fields are left unknown rather than estimated."},{"date":"2026-07-25","tag":"model","text":"Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive thinking by default, $5/$25 per MTok base pricing, and official cloud-platform availability. Added independent Artificial Analysis evidence (61 Intelligence Index at max effort; 52.3 output tok/s) with effort-specific caveats. Gemini 3.5 Flash Cyber remains limited to CodeMender government and trusted-partner pilots; GPT-Live and Muse Spark 1.1 remain non-API products, so none were added to the API roster. Claude Opus 4.7 Fast Mode was removed July 24."},{"date":"2026-07-24","tag":"benchmark","text":"Refreshed current-source evidence: added OpenAI's GPT-5.6 launch table, Google's Managed Agents update, and Scale's public SWE-bench Pro leaderboard. Clarified that vendor launch tables and the public leaderboard are not directly comparable because their model versions and harnesses differ."},{"date":"2026-07-22","tag":"fix","text":"Made the selected reference propagate through open-weight quality headlines, market-signal history, comparison headings, model analytics, self-hosting quality and capability market position. Replaced the model and harness inventory mini-lines with stacked proprietary/open category areas based on release-tag roster snapshots."},{"date":"2026-07-22","tag":"model","text":"Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M context, 64K output, current API pricing, Artificial Analysis intelligence and throughput measurements, and Google’s published coding and agentic benchmarks."},{"date":"2026-07-22","tag":"data","text":"Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to governments and trusted CodeMender partners; Gemini Omni Flash and Nano Banana 2 Lite remain specialized media models rather than general-purpose roster entries."},{"date":"2026-07-21","tag":"model","text":"Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint: fully open Apache 2.0 weights and data, 4K context, 1.7T-token distillation, 1,811 languages, and official BF16, FP8, NVFP4A16, INT3, INT4 and INT6 variants."},{"date":"2026-07-21","tag":"harness","text":"Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI 0.144.6, Gemini CLI 0.51.0, OpenCode 1.18.4, Cline 4.0.10, and Goose 1.43.0; updated repository star snapshots and notable release capabilities."},{"date":"2026-07-21","tag":"model","text":"Added Moonshot's official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached input, and $15/M output; the 30/70 workload blend is $11.40/M before reasoning-effort effects."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark registers and selectable reference configurations. Added documented quality, speed, cost and capability composites in data/report-metrics.json."},{"date":"2026-07-18","tag":"routing","text":"Updated the action queue and recommended routing: Fable for hardest retained-data workloads, Terra for default engineering, Luna for high-volume subagents, with explicit escalation rules."},{"date":"2026-06-06","tag":"fix","text":"Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K context, 1000+ tok/s on Cerebras) and Hermes Agent v0.16.0 (Nous Research, MIT, self-hosted multi-platform agent) — both were incorrectly removed in v1.5.0 sweep."},{"date":"2026-06-06","tag":"policy","text":"Dashboard market sweep v1.5.0: real-world re-grounding. Replaced fictional Mythos/GPT-5.5-Cyber rows with verified models; added Nvidia Nemotron coalition, Kimi K2.6, GLM-5, Cohere Command A+, SubQ 1M-Preview."},{"date":"2026-06-04","tag":"model","text":"Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context, NVIDIA Open Model License) at Computex — first frontier-scale open model from Nvidia."},{"date":"2026-06-04","tag":"model","text":"Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral, Perplexity, Reflection AI, Sarvam, Thinking Machines Lab as inaugural members."},{"date":"2026-06-01","tag":"model","text":"Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model."}],"actions":["P0 · Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","P0 · Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","P0 · Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","P1 · Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","P1 · Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","P2 · Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics."],"models":[{"id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","tier":"frontier","released":"2026-06-09","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":80,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii":59.9,"coding_agent_index":77.2,"deep_swe":69.7,"terminal_bench":83.1,"agents_last_exam":40.5,"api_in":10,"api_out":50,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":52.3,"subscription":"Claude API and supported cloud platforms","notes":"Default reference for v2. Anthropic describes Fable 5 as its most capable widely released model. Adaptive thinking is always on. Comparative latency is slower. Retention is 30 days and zero-data-retention is not available.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"status":"stable","speed_class":"deliberate","speed_evidence":"vendor-qualitative","capability_levels":{"coding":49,"reasoning":50,"knowledge":48,"comms":48,"multimodal":45,"agentic":50}},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":64.6,"swe_verified":null,"aaii":58.9,"coding_agent_index":80,"deep_swe":72.7,"terminal_bench":88.8,"agents_last_exam":52.7,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Flagship GPT-5.6 tier. API Fast mode replaced Priority Processing on July 30: OpenAI claims up to 2.5× Standard speed at 2× Standard price with no intelligence change. Speed is otherwise stored as an end-to-end task index in report-metrics.json, not tok/s.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":49,"knowledge":48,"comms":48,"multimodal":48,"agentic":49}},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":63.4,"swe_verified":null,"aaii":55,"coding_agent_index":77.4,"deep_swe":69.6,"terminal_bench":87.4,"agents_last_exam":50.4,"api_in":2,"api_out":12,"api_cache_hit":0.2,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Balanced GPT-5.6 tier and recommended default engineering route. On July 30, API list pricing fell 20% from $2.50/$15 to $2/$12 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":48,"reasoning":46,"knowledge":46,"comms":46,"multimodal":45,"agentic":48}},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","tier":"fast","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":62.7,"swe_verified":null,"aaii":51.2,"coding_agent_index":74.6,"deep_swe":67.2,"terminal_bench":84.7,"agents_last_exam":50.3,"api_in":0.2,"api_out":1.2,"api_cache_hit":0.02,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Fastest, most affordable GPT-5.6 tier; recommended for high-volume subagents with verification and escalation. On July 30, API list pricing fell 80% from $1/$6 to $0.20/$1.20 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change. ChatGPT is also rolling Luna out as the Free and Go default, with unlimited text chats and a Think option; that product change does not alter API routing.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":46,"reasoning":44,"knowledge":43,"comms":45,"multimodal":43,"agentic":46}},{"id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","tier":"frontier","released":"2026-07-16","license":"open weights announced; terms pending weight release","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":null,"tok_per_sec":null,"subscription":"Kimi API, Kimi Code, Kimi Work","reasoning_capable":true,"effort_default":"max","effort_levels":["max"],"notes":"2.8T sparse MoE with 16/896 experts active, native vision and always-on thinking. API pricing is $3/M uncached input, $0.30/M cached input and $15/M output. Full weights promised by July 27, 2026.","quality_vs_fable":99.2,"quality_evidence":"Geometric mean across 14 overlapping values in Moonshot’s launch comparison; vendor-reported, max/xhigh settings.","capability_levels":{"coding":47,"reasoning":46,"knowledge":47,"comms":46,"multimodal":47,"agentic":48}},{"id":"opus-5","name":"Claude Opus 5","provider":"Anthropic","tier":"frontier","released":"2026-07-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"Claude API, Bedrock, Google Cloud, Microsoft Foundry","notes":"Current Anthropic Opus generation. 1M context and 128K maximum output. Adaptive thinking is enabled by default; effort defaults to high. Artificial Analysis reports a 61 Intelligence Index and 52.3 output tok/s at max effort; these figures are effort-specific. Research-preview Fast Mode is Claude API-only at $10/$50 per MTok and claims up to 2.5x higher output throughput.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"speed_class":"balanced","speed_evidence":"vendor-qualitative","cite":[86,87]},{"id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","tier":"frontier","released":"2026-05-28","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":69.2,"swe_verified":88.6,"livecodebench":82,"aime":90,"tau2":86,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Max 20× $200/mo · Max 5× $100/mo","notes":"Previous Opus generation, retained as an active compatibility option. SWE-V 88.6%, SWE-Pro 69.2%, AAII 61.4. Fast Mode reduced 3× to $10/$50 (was $30/$150 on 4.7). 1M context standard.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":50,"comms":50,"multimodal":34,"agentic":50}},{"id":"opus-4.7","name":"Claude Opus 4.7","provider":"Anthropic","tier":"frontier","released":"2026-04-16","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":64.3,"swe_verified":87.6,"livecodebench":79,"aime":88,"tau2":84,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":50,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Now legacy as of Opus 4.8 release May 28. Pricing unchanged. SWE-Verified 87.6%, SWE-Pro 64.3%. Fast Mode was removed July 24, 2026; standard-speed API access remains active.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":49,"comms":50,"multimodal":32,"agentic":50}},{"id":"sonnet-4.6","name":"Claude Sonnet 4.6","provider":"Anthropic","tier":"frontier","released":"2026-02-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":44,"swe_verified":77,"livecodebench":76,"aime":85,"tau2":81,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":50,"tok_per_sec":80,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Best code style/intent understanding. With cache+batch: $0.30/$7.50 effective.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","max"],"cache_discount":0.1,"capability_levels":{"coding":41,"reasoning":40,"knowledge":39,"comms":42,"multimodal":31,"agentic":41}},{"id":"haiku-4.5","name":"Claude Haiku 4.5","provider":"Anthropic","tier":"fast","released":"2025-10-15","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":28,"swe_verified":62,"livecodebench":58,"aime":70,"tau2":65,"api_in":1,"api_out":5,"api_cache_hit":0.1,"batch_discount":50,"tok_per_sec":110,"subscription":"Available in all tiers","notes":"5× cheaper than Sonnet. Triage/classification champion.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.1,"capability_levels":{"coding":28,"reasoning":18,"knowledge":29,"comms":30,"multimodal":18,"agentic":28}},{"id":"gpt-5.5","name":"GPT-5.5","provider":"OpenAI","tier":"frontier","released":"2026-04-23","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M (272K threshold)","swe_pro":58.6,"swe_verified":88.7,"livecodebench":84,"aime":92,"tau2":82,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Pro $200 · Pro Lite $100 · Plus $20","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"OpenAI flagship; default in ChatGPT (Instant variant since May 5, 2026). 1M context with 2× input/1.5× output surcharge above 272K. Reasoning tokens billed as output.","cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":41,"knowledge":49,"comms":40,"multimodal":41,"agentic":38}},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","provider":"OpenAI","tier":"frontier","released":"2026-04-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":60,"swe_verified":90,"livecodebench":86,"aime":94,"tau2":84,"api_in":30,"api_out":180,"api_cache_hit":3,"batch_discount":50,"tok_per_sec":35,"subscription":"Pro $200 only","reasoning_capable":true,"effort_default":"high","effort_levels":["medium","high","xhigh"],"notes":"Highest-stakes reasoning tier. $30/$180. Available in ChatGPT Pro $200 and as API. 6× cost of base 5.5.","cache_discount":0.25,"capability_levels":{"coding":47,"reasoning":48,"knowledge":50,"comms":42,"multimodal":42,"agentic":42}},{"id":"gpt-5.4","name":"GPT-5.4","provider":"OpenAI","tier":"frontier","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":57.7,"swe_verified":81,"livecodebench":81,"aime":90,"tau2":80,"api_in":2.5,"api_out":15,"api_cache_hit":0.25,"batch_discount":50,"tok_per_sec":60,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Best quality-per-credit. Held SWE-Pro lead Feb-April. Default for everyday coding.","cache_discount":0.1,"capability_levels":{"coding":40,"reasoning":40,"knowledge":40,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-mini","name":"GPT-5.4 Mini","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":38,"swe_verified":73,"livecodebench":71,"aime":78,"tau2":70,"api_in":0.4,"api_out":1.6,"api_cache_hit":0.04,"batch_discount":50,"tok_per_sec":130,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"94% of GPT-5.4 coding at 6× less. Best subagent. ~1/20 burn vs GPT-5.5. Collapses at 64K+ context.","cache_discount":0.1,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-nano","name":"GPT-5.4 Nano","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":25,"swe_verified":60,"livecodebench":58,"aime":65,"tau2":55,"api_in":0.1,"api_out":0.4,"api_cache_hit":0.01,"batch_discount":50,"tok_per_sec":200,"subscription":"API only","reasoning_capable":true,"effort_default":"low","effort_levels":["minimal","low","medium"],"notes":"Smallest reasoning model. API-only. For embeddable/edge inference at near-zero cost.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":20,"comms":20,"multimodal":20,"agentic":20}},{"id":"gpt-5.3-codex","name":"GPT-5.3-Codex","provider":"OpenAI","tier":"frontier","released":"2026-01-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":56.8,"swe_verified":85,"livecodebench":82,"aime":87,"tau2":78,"api_in":1.5,"api_out":10,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":70,"subscription":"All paid tiers · Code Review uses this","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Coding-specialised. ~⅓ burn vs GPT-5.5 for ~2pts less SWE-Pro. Often best $/quality for execution turns.","cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":28,"knowledge":28,"comms":27,"multimodal":8,"agentic":38}},{"id":"gpt-5.3-codex-spark","name":"GPT-5.3-Codex-Spark","provider":"OpenAI","tier":"fast","released":"2026-02-12","license":"proprietary","jurisdiction":"US","context":128000,"context_label":"128K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":1000,"subscription":"ChatGPT Pro — research preview only","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"Smaller, latency-first sibling of GPT-5.3-Codex. Released Feb 12, 2026. 1000+ tok/s on Cerebras hardware. ChatGPT Pro research preview only — not in API at launch; separate preview rate-limit pool (no standard credit burn). Text-only. Target use: real-time micro-edits, live pair-programming in Codex app/CLI/VS Code.","cache_discount":null,"capability_levels":{"coding":38,"reasoning":22,"knowledge":22,"comms":25,"multimodal":0,"agentic":28}},{"id":"gpt-5.2","name":"GPT-5.2","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":52,"swe_verified":78,"livecodebench":76,"aime":84,"tau2":75,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Available but not recommended","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Codex picker keeps it for 'long-running agents' (specifically tuned for autonomy). Otherwise eclipsed by 5.3-Codex.","cache_discount":0.1},{"id":"gpt-5.2-codex","name":"GPT-5.2-Codex","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":50,"swe_verified":76,"livecodebench":75,"aime":82,"tau2":73,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Legacy","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Predecessor to 5.3-Codex. Legacy.","cache_discount":0.1},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","provider":"OpenAI","tier":"fast","released":"2026-05-05","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":35,"swe_verified":76,"livecodebench":72,"aime":81,"tau2":70,"api_in":1.5,"api_out":6,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":200,"subscription":"Default ChatGPT model · API chat-latest","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"ChatGPT default since May 5, 2026. 52.5% fewer hallucinations vs 5.3, 30% shorter responses, first Instant-class High-capability rating on cybersec/bio-chem.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":40,"comms":30,"multimodal":30,"agentic":10}},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","provider":"Google","tier":"frontier","released":"2026-02-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":54.2,"swe_verified":78,"livecodebench":76,"aime":85,"tau2":78,"api_in":2,"api_out":12,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":119,"subscription":"Gemini Advanced $20 · Ultra $100","notes":"Released Feb 19, 2026 (preview). Pricing doubles above 200K input tokens. 50% batch discount.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":39,"reasoning":41,"knowledge":49,"comms":41,"multimodal":50,"agentic":29}},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":58.7,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78,"agents_last_exam":null,"quality_vs_fable":81.6,"api_in":1.5,"api_out":7.5,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":304,"subscription":"Gemini API · AI Studio · Gemini app · Antigravity","notes":"GA stable ID gemini-3.6-flash. Google reports fewer tool calls and 17% fewer output tokens than 3.5 Flash on the AA Index workload; Computer Use is preview. Artificial Analysis measured about 304 output tok/s at high thinking.","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":42,"reasoning":38,"knowledge":45,"comms":40,"multimodal":47,"agentic":42}},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","provider":"Google","tier":"legacy","released":"2026-05-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55.1,"swe_verified":78,"livecodebench":70,"aime":75,"tau2":68,"aaii_v4_1":50,"swe_bench_pro":55.1,"deep_swe_v1_1":37,"terminal_bench_2_1":76.2,"agents_last_exam":null,"quality_vs_fable":76.1,"api_in":1.5,"api_out":9,"api_cache_hit":0.375,"batch_discount":50,"tok_per_sec":165,"subscription":"Free CLI: 1000 req/day at 1M context","notes":"Shipped GA at Google I/O May 19, 2026. Still offered, but Gemini 3.6 Flash is the recommended migration target with stronger agentic results and lower output-token pricing.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":30,"reasoning":30,"knowledge":40,"comms":30,"multimodal":40,"agentic":20}},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":54.2,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":36,"swe_bench_pro":54.2,"deep_swe_v1_1":null,"terminal_bench_2_1":54,"agents_last_exam":null,"quality_vs_fable":65.9,"api_in":0.3,"api_out":2.5,"api_cache_hit":0.03,"batch_discount":50,"tok_per_sec":490,"subscription":"Gemini API · AI Studio · Gemini app rollout","notes":"GA stable ID gemini-3.5-flash-lite. Google positions it for high-volume subagents, document parsing and structured extraction; Artificial Analysis measured about 490 output tok/s. Computer Use availability differs across Google documentation and should be validated per API surface.","reasoning_capable":true,"effort_default":"minimal","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":32,"reasoning":28,"knowledge":34,"comms":32,"multimodal":38,"agentic":35}},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash Lite","provider":"Google","tier":"legacy","released":"2026-03-10","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":22,"swe_verified":65,"livecodebench":60,"aime":68,"tau2":58,"api_in":0.25,"api_out":1.5,"api_cache_hit":0.0625,"batch_discount":50,"tok_per_sec":220,"subscription":"Free CLI + AI Studio","notes":"Lite tier; 1M context retained.","reasoning_capable":true,"effort_default":"off","effort_levels":["off","low","medium"],"cache_discount":0.25,"capability_levels":{"coding":22,"reasoning":22,"knowledge":30,"comms":25,"multimodal":35,"agentic":15}},{"id":"deepseek-v4-pro","name":"DeepSeek V4-Pro","provider":"DeepSeek","tier":"frontier","released":"2026-04-24","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":55.4,"swe_verified":80.6,"livecodebench":93.5,"aime":88,"tau2":78,"api_in":0.435,"api_out":0.87,"api_cache_hit":0.043,"batch_discount":null,"tok_per_sec":60,"subscription":"API only","notes":"MIT-licensed. Permanent pricing May 22, 2026 — Opus-class quality at ~1/10 cost. 1.6T/49B MoE, 1M context.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.0083,"capability_levels":{"coding":47,"reasoning":40,"knowledge":38,"comms":30,"multimodal":8,"agentic":28}},{"id":"deepseek-v4-flash","name":"DeepSeek V4-Flash","provider":"DeepSeek","tier":"fast","released":"2026-07-31","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"tau2":72,"terminal_bench_2_1":82.7,"agents_last_exam":25.2,"api_in":0.14,"api_out":0.28,"api_cache_hit":0.0028,"batch_discount":null,"tok_per_sec":90,"subscription":"API + open weights","notes":"V4-Flash-0731 public beta; 284B/13B active MoE, 1M context and 384K max output. DeepSeek reports Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in its unreleased minimal harness; internal DSBench results are not normalized here. A future 2x peak-hours rate is announced, but no effective date is published.","reasoning_capable":true,"effort_default":"think","effort_levels":["non-think","think"],"cache_discount":0.02,"capability_levels":{"coding":40,"reasoning":30,"knowledge":30,"comms":25,"multimodal":5,"agentic":25}},{"id":"minimax-m2.7","name":"MiniMax M2.7","provider":"MiniMax","tier":"frontier","released":"2026-03-18","license":"open-weight","jurisdiction":"China","context":205000,"context_label":"205K","swe_pro":40,"swe_verified":74,"livecodebench":72,"aime":80,"tau2":73,"api_in":0.3,"api_out":1.2,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":70,"subscription":"API + open-weight","notes":"Current flagship reasoner; MoE 230B/10B active.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.52,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":20,"agentic":20}},{"id":"grok-4.3","name":"Grok 4.3","provider":"xAI","tier":"frontier","released":"2026-05-06","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":45,"swe_verified":80,"livecodebench":80,"aime":88,"tau2":76,"api_in":1.25,"api_out":2.5,"api_cache_hit":0.31,"batch_discount":null,"tok_per_sec":75,"subscription":"SuperGrok $30 · Heavy $300","notes":"Current xAI flagship — aggressive pricing for frontier tier. Hybrid reasoning. Multimodal text+image.","reasoning_capable":true,"effort_default":"reasoning-on","effort_levels":["off","on"],"cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":44,"knowledge":42,"comms":34,"multimodal":34,"agentic":34}},{"id":"grok-4.1-fast","name":"Grok 4.1 Fast","provider":"xAI","tier":"fast","released":"2026-03-20","license":"proprietary","jurisdiction":"US","context":2000000,"context_label":"2M","swe_pro":30,"swe_verified":70,"livecodebench":65,"aime":75,"tau2":65,"api_in":0.2,"api_out":0.5,"api_cache_hit":0.05,"batch_discount":null,"tok_per_sec":140,"subscription":"SuperGrok $30","notes":"Cheapest large-context model on market. 2M context, $0.05/Mtok cached input.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.25,"capability_levels":{"coding":28,"reasoning":28,"knowledge":30,"comms":26,"multimodal":26,"agentic":22}},{"id":"mistral-medium-3.5","name":"Mistral Medium 3.5","provider":"Mistral","tier":"frontier","released":"2026-04-29","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":42,"swe_verified":77.6,"livecodebench":75,"aime":75,"tau2":65,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":75,"subscription":"Le Chat Pro €20","notes":"EU jurisdiction. 128B dense, Apache 2.0. Strongest non-Chinese open-weight coding agent. Vibe agents (GitHub/Linear/Jira/Sentry integrations). Pricing not verified; Medium 3 was $0.40/$2.00.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":40,"reasoning":30,"knowledge":40,"comms":30,"multimodal":10,"agentic":20}},{"id":"mistral-large-3","name":"Mistral Large 3","provider":"Mistral","tier":"frontier","released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"tau2":70,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":65,"subscription":"API + Le Chat","notes":"Cheapest premium output price in market among Western frontier. 675B/41B MoE, Apache 2.0.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":38,"reasoning":36,"knowledge":40,"comms":36,"multimodal":20,"agentic":28}},{"id":"subq-1m-preview","name":"SubQ 1M-Preview","provider":"SubQ","tier":"frontier","released":"2026-05-15","license":"proprietary","jurisdiction":"US","context":12000000,"context_label":"12M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":1,"api_out":5,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":80,"subscription":"API preview","notes":"First commercial subquadratic (non-transformer) LLM. ~1/5 frontier cost on long-context tasks. Capability rating estimated — public benchmarks pending.","reasoning_capable":null,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":30,"reasoning":32,"knowledge":36,"comms":30,"multimodal":5,"agentic":28}},{"id":"nemotron-3-ultra","name":"Nvidia Nemotron 3 Ultra","provider":"Nvidia","tier":"frontier","released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"tau2":76,"api_in":1.5,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":300,"subscription":"build.nvidia.com (closed API tier) + open weights","notes":"550B/55B MoE, hybrid Mamba-Transformer. ~300 tok/s. Open weights also available — see self_hosting. Capability levels are best estimates pending independent benchmarks.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":42,"reasoning":42,"knowledge":42,"comms":34,"multimodal":10,"agentic":34}},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","provider":"Swiss AI Initiative","tier":"fast","released":"2026-06-15","license":"Apache 2.0","jurisdiction":"Switzerland","context":4096,"context_label":"4K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Self-hosted weights","notes":"Largest newly released Apertus v1.1 distilled checkpoint. Dense 4.6B storage / 3.8B compute parameters, trained on 1.7T tokens, supports 1,811 languages, and ships in BF16 plus server and Apple-oriented quantizations. No comparable coding-agent, throughput, or API-price evidence is published.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null},{"id":"inkling","name":"Inkling","provider":"Thinking Machines Lab","tier":"frontier","released":"2026-07-15","license":"Apache 2.0","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Open weights (Hugging Face); fine-tuning and inference via the Tinker platform and third-party providers","notes":"975B-parameter multimodal MoE (41B active), 66-layer decoder-only transformer, 1M-token context. Text/image/audio input, text-only output. No first-party per-token API pricing published -- Thinking Machines monetizes via the Tinker fine-tuning platform rather than metered inference. A smaller Inkling-Small (12B active) companion model was released alongside it. Benchmark results are published on the model card across reasoning, agentic, coding, factuality, vision, audio, and safety categories but are not yet normalized into this roster's comparable benchmark set.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null}],"subscriptions":[{"provider":"Anthropic","tier":"Pro","price_usd":20,"limits":"~45 messages / 5h on Sonnet · limited Opus","models":"Sonnet 4.6, Haiku 4.5, limited Opus 4.8/4.7","features":"Chat only (Claude Code removed April 2026). From June 15, 2026 split into Chat pool + Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 5×","price_usd":100,"limits":"5× Pro quotas · ~225 msg/5h Sonnet · expanded Opus","models":"Full Opus 4.8/4.7 · Sonnet 4.6 · Haiku 4.5","features":"Cache reads included flat-rate. From June 15, 2026: Chat pool + separate Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 20×","price_usd":200,"limits":"20× Pro quotas","models":"All","features":"For heavy Opus users. From June 15, 2026: Chat + Agent SDK credit pools."},{"provider":"OpenAI","tier":"Plus","price_usd":20,"limits":"~80 GPT-5.4 msg/3h","models":"GPT-5.4 (limited) · GPT-5.4 Mini · o-series","features":"ChatGPT · GPTs · Codex CLI 30-150 tasks/5h"},{"provider":"OpenAI","tier":"Pro 5×","price_usd":100,"limits":"5× Plus quotas (new tier April 2026)","models":"GPT-5.4 Thinking unlimited","features":"Released as Anthropic Max competitor"},{"provider":"OpenAI","tier":"Pro","price_usd":200,"limits":"Effectively unlimited","models":"All including o3-pro","features":"Original premium tier"},{"provider":"Google","tier":"Gemini Advanced","price_usd":20,"limits":"Generous, soft caps","models":"Gemini 3.1 Pro · Gemini 3.6 Flash · Gemini 3.5 Flash-Lite","features":"Workspace integration · 1M context"},{"provider":"Google","tier":"Ultra","price_usd":100,"limits":"Higher quotas + Veo video","models":"All Gemini + research preview","features":"Veo 3 video · Project Mariner"},{"provider":"Mistral","tier":"Le Chat Pro","price_usd":22,"limits":"Generous","models":"Mistral Large 3 · Codestral","features":"EU jurisdiction"},{"provider":"xAI","tier":"SuperGrok","price_usd":30,"limits":"Generous","models":"Grok 4 · Grok 4 Heavy","features":"X integration"},{"provider":"DeepSeek","tier":"API only","price_usd":null,"limits":"Pay per token","models":"V4-Pro · V4-Flash","features":"Cheapest frontier API ($0.435/$0.87 permanent since May 22, 2026)"}],"agent_policies":[{"provider":"Anthropic","subscription_automated":"Prohibited","enforcement":"Active (OAuth blocked Apr 4)","first_party_exception":"`claude -p` pipe mode and Claude Code itself","api_required_for_automation":true,"cite":[35]},{"provider":"OpenAI","subscription_automated":"Prohibited (ToS)","enforcement":"Currently tolerated","first_party_exception":"Codex CLI uses subscription quota","api_required_for_automation":false},{"provider":"Google","subscription_automated":"Allowed in CLI","enforcement":"—","first_party_exception":"Gemini CLI free tier 1000 req/day","api_required_for_automation":false},{"provider":"Mistral","subscription_automated":"Allowed","enforcement":"—","first_party_exception":"—","api_required_for_automation":false}],"harnesses":[{"id":"claude-code","name":"Claude Code","vendor":"Anthropic","license":"proprietary","category":"CLI + IDE","stars":49000,"providers":["Anthropic"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":true,"computer_use":true,"lsp":true,"git":true,"memory":"CLAUDE.md","sandbox":"local","swe_pro":46,"pricing":"Subscription Pro/Max","sweet_spot":"SubagentStop hooks, /plugin list, requiredMinimumVersion managed setting, MCP fixes, Agent Teams.","stumbles":"Anthropic-only. Removed from standard Pro tier April 2026 — push to Max.","cite":[23,1]},{"id":"opencode","name":"OpenCode","vendor":"Anomaly","license":"MIT","category":"CLI + ACP","stars":188119,"providers":["75+: Anthropic*, OpenAI, Google, Mistral, Kimi, GLM, Ollama, LM Studio"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.18.4 (Jul 20, 2026): adaptive Kimi thinking controls, provider-defined reasoning options, restored Azure endpoints, and a rewritten desktop prompt input.","stumbles":"Anthropic OAuth blocked April 4 — must use API key.","cite":[24,35,72]},{"id":"codex-cli","name":"Codex CLI","vendor":"OpenAI","license":"Apache 2.0","category":"CLI + macOS app","stars":100232,"providers":["OpenAI"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"cloud","swe_pro":56.8,"pricing":"Plus $20 / Pro $200","sweet_spot":"v0.144.6 (Jul 18, 2026): refreshed GPT-5.6 Sol/Terra/Luna bundled instructions and corrected Codex context-window metadata to 272K.","stumbles":"OpenAI-only. No MCP, no hooks. Tightly coupled to apply_patch tool.","cite":[25,36,70]},{"id":"gemini-cli","name":"Gemini CLI","vendor":"Google","license":"Apache 2.0","category":"CLI","stars":106096,"providers":["Google"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"GEMINI.md","sandbox":"local","swe_pro":null,"pricing":"Free 1000 req/day","sweet_spot":"v0.51.0 (Jul 16, 2026): hardened sensitive-path and symlink handling, read-only macOS sandbox git config, and modern-model escape-sequence fixes.","stumbles":"Sunsetting to Antigravity CLI for free tier on June 18, 2026; paid Gemini/Enterprise keys retain access.","cite":[26,71]},{"id":"aider","name":"Aider","vendor":"paul-gauthier","license":"Apache 2.0","category":"CLI","stars":32000,"providers":["Anthropic*","OpenAI","Google","Ollama","100+"],"mcp":false,"skills":false,"hooks":false,"subagents":false,"voice":true,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"CONVENTIONS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"Mature, lightweight, voice-input native. Pair-programmer mode. Active, last commit Mar 2026. 44k stars.","stumbles":"No MCP, no hooks. Less ambitious than newer harnesses.","cite":[27]},{"id":"cline","name":"Cline","vendor":"cline-bot","license":"Apache 2.0","category":"VS Code extension","stars":64886,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".clinerules","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v4.0.10 (Jul 20, 2026): current release adds telemetry for consecutive-mistake-limit events; multi-editor and CLI surfaces remain available.","stumbles":"Anthropic OAuth blocked. Can be expensive on long sessions.","cite":[28,35,73]},{"id":"roo-code","name":"Roo Code","vendor":"RooVetGit","license":"Apache 2.0","category":"VS Code extension","stars":18000,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".roo","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v3.53.0 (Apr 23, 2026). Power-user Cline fork, model-agnostic, BYOK. Recent: GPT-5.5 via Codex, Opus 4.7 on Vertex, checkpoint nav.","stumbles":"Same OAuth situation. Configuration complexity.","cite":[29,35]},{"id":"cursor","name":"Cursor","vendor":"Anysphere","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".cursorrules","sandbox":"local","swe_pro":null,"pricing":"Hobby free · Pro $20 · Pro+ $60 · Ultra $200 · Teams $40/user","sweet_spot":"Cursor 3.5 (May 20, 2026): Cloud Agents (isolated VMs, multi-repo), Composer 2.5, Agents Window, parallel subagents.","stumbles":"Closed source. Lock-in. Subscription required for serious use.","cite":[31]},{"id":"windsurf","name":"Windsurf / Devin Desktop","vendor":"Cognition","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"windsurfrules","sandbox":"local","swe_pro":null,"pricing":"Pro $20 · Max $200","sweet_spot":"Renamed to Devin Desktop, Agent Command Center kanban, embedded Devin cloud agent, SWE-1.6 model, multi-agent + worktrees.","stumbles":"Pro $20 (was $15); smaller ecosystem than Cursor.","cite":[32]},{"id":"goose","name":"Goose","vendor":"Block","license":"Apache 2.0","category":"Desktop + CLI","stars":51387,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":false,"git":true,"memory":".goosehints","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.43.0 (Jul 14, 2026): per-message token/cost/TTFT/tok-s metrics, ACP reconnection, GPT-5.6 support, dynamic Ollama Cloud discovery, and expanded providers.","stumbles":"Anthropic OAuth blocked. Less mindshare than OpenCode.","cite":[30,35,74]},{"id":"omo","name":"OMO (Multi-model orchestrator)","vendor":"community","license":"MIT","category":"Multi-agent orchestrator","stars":54000,"providers":["multi-provider"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free","sweet_spot":"Rebrand from oh-my-opencode; multi-model orchestration. Wraps Claude Code, OpenCode, Codex, Kimi K2, DeepSeek V4, Gemini CLI.","stumbles":"Niche. Steep learning curve."},{"id":"hermes","name":"Hermes Agent","vendor":"Nous Research","license":"MIT","category":"Self-hosted multi-platform agent","stars":null,"providers":["OpenRouter-style multi-model"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":true,"remote":true,"computer_use":true,"lsp":false,"git":true,"memory":"persistent memory + auto-gen skills","sandbox":"local/docker/ssh/singularity/modal","swe_pro":null,"pricing":"Free (self-hosted)","sweet_spot":"v0.16.0. Persistent memory + auto-generated skills — learns your projects. Bridges Telegram/Discord/Slack/WhatsApp/Signal/Email/CLI. Natural-language cron for unattended runs. Parallel isolated subagents. Web search, browser automation, vision, image-gen, TTS.","stumbles":"Not coding-specialised — general autonomous agent. Setup requires self-host script. No first-party SWE benchmark.","cite":[]},{"id":"github-copilot-cli","name":"GitHub Copilot CLI","vendor":"GitHub","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["GitHub Models"],"mcp":false,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":true,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Bundled with Copilot Business/Enterprise","sweet_spot":"GA Feb 25, 2026. Specialized sub-agents (Explore, Task, Code Review, Plan), background delegation, autopilot.","stumbles":"Bundled-only — no standalone tier. GitHub-centric."},{"id":"amp-cli","name":"Amp CLI","vendor":"Sourcegraph","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["Multi"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Subscription (Sourcegraph)","sweet_spot":"Spun out as standalone company 2026. Runs as sidebar agent inside Zed via Terminal Threads.","stumbles":"Early standalone phase."},{"id":"zed","name":"Zed","vendor":"Zed Industries","license":"proprietary","category":"IDE","stars":null,"providers":["15 LLM providers + MCP"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Personal free (2k predictions) · Pro $10 · Business $30/seat","sweet_spot":"Rust-native editor with first-class agent panel + ACP host. Terminal Threads run Claude Code/Amp inline.","stumbles":"Editor first; agent layer still maturing."},{"id":"continue-dev","name":"Continue","vendor":"Continue","license":"Apache 2.0","category":"IDE extension + CLI","stars":null,"providers":["multi-provider"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".continue","sandbox":"local","swe_pro":null,"pricing":"Solo $0 · Team/Company ~$10/dev/mo","sweet_spot":"Agent mode plan+execute. Continuous AI, Mission Control, shared PR/ticket workflows.","stumbles":"Newer agent features still stabilizing across providers."}],"self_hosting":{"hardware_options":[{"id":"vast-2xa6000","name":"Vast.ai 2× A6000","vram_gb":96,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Reference cloud-GPU setup for this dashboard. No upfront capex. Docker templates, SSH/Cloudflare Zero Trust. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-3xa6000","name":"Vast.ai 3× A6000","vram_gb":144,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Headroom for Qwen 3 235B-A22B Q6_K. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-4xa6000","name":"Vast.ai 4× A6000","vram_gb":192,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Llama 3.1 405B Q4 territory. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"mbp-14-m4pro-64","name":"MacBook Pro 14\" M4 Pro 64GB","vram_gb":64,"cost_label":"~$3,200 capex","cost_per_hour":null,"type":"local","notes":"MoE sweet spot. 273 GB/s memory bandwidth."},{"id":"mbp-16-m5max-128","name":"MacBook Pro 16\" M5 Max 128GB","vram_gb":128,"cost_label":"~$5,500 capex","cost_per_hour":null,"type":"local","notes":"Best portable inference. ~545 GB/s bandwidth."},{"id":"mba-15-32","name":"MacBook Air 15\" 32GB","vram_gb":32,"cost_label":"~$1,900 capex","cost_per_hour":null,"type":"local","notes":"Hard 32GB ceiling. Limited to ~14B dense or 26B MoE Q4."},{"id":"contabo-2xl40s","name":"Contabo 2 x L40S","provider":"Contabo","vram_gb":96,"cost_label":"€1,502/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"EU-oriented fixed monthly configuration; 64 vCPU, 213 GB RAM, 3.5 TB storage, and 15 TB bandwidth are listed for the 2-GPU tier."},{"id":"contabo-1xh200","name":"Contabo 1 x H200 NVL","provider":"Contabo","vram_gb":141,"cost_label":"€2,149/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"Fixed monthly large-memory option. Confirm location and availability before treating it as a sovereignty or latency fit."},{"id":"infomaniak-1xl40s","name":"Infomaniak 1 x L40S","provider":"Infomaniak","vram_gb":48,"cost_label":"Live calculator / availability validation","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"calculator_or_request","price_checked":"2026-07-19","source":"https://www.infomaniak.com/en/hosting/public-cloud/prices","notes":"Swiss OpenStack option with dedicated GPU access and usage billing. Public pages list L40S availability but do not expose a stable crawlable SKU price; validate availability for the selected region."},{"id":"hyperstack-1xh200","name":"Hyperstack 1 x H200 SXM","provider":"Hyperstack","vram_gb":141,"cost_label":"$3.50/hr on demand","cost_per_hour":3.5,"type":"cloud","show_in_fit":false,"price_status":"published_on_demand","price_checked":"2026-07-19","source":"https://www.hyperstack.cloud/","notes":"Minute-accurate on-demand billing; reservation pricing starts at $2.45/hr. Validate region, storage, and availability."}],"models":[{"id":"gemma-4-26b-moe","name":"Gemma 4 26B-A4B MoE","params_total":26,"params_active":4,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":35,"livecodebench":77,"aime":88,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-3xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-4xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":22,"tok_per_sec":75},"mbp-16-m5max-128":{"quant":"BF16","vram_used":52,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"MoE — only 4B active per token. Faster than dense models 5× its size."},{"id":"gemma-4-31b-dense","name":"Gemma 4 31B Dense","params_total":31,"params_active":31,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":38,"livecodebench":80,"aime":89,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-3xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-4xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":19,"tok_per_sec":28},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":26,"tok_per_sec":38},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Highest quality open Gemma. Slower per-token (full 31B active)."},{"id":"qwen-3.6-plus","name":"Qwen 3.6 Plus","params_total":397,"params_active":17,"released":"2026-04-11","license":"Apache 2.0","jurisdiction":"China","swe_pro":50,"livecodebench":71.4,"aime":87,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":38},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":35},"vast-4xa6000":{"quant":"Q8_0","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1M context. Top open agentic coder. April 11 release.","swe_verified":68.2},{"id":"llama-4-maverick","name":"Llama 4 Maverick","params_total":400,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":42,"livecodebench":70,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":28},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":25},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"400B / 17B active MoE. 1M context. Strong MMLU-Pro (80.5%) but coding behind Chinese labs.","swe_verified":72},{"id":"minimax-m2.5-open","name":"MiniMax M2.5 (open weights)","params_total":456,"params_active":46,"released":"2026-01-20","license":"open-weight","jurisdiction":"China","swe_pro":40,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Eclipsed by DeepSeek V4 and Qwen 3.6 Plus. China jurisdiction. Maintained for niche workloads."},{"id":"deepseek-v4-pro-open","name":"DeepSeek V4-Pro (open weights)","params_total":1600,"params_active":49,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":55.4,"livecodebench":93.5,"aime":90,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q2_K","vram_used":188,"tok_per_sec":14},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1.6T total / 49B active MoE. MIT license. Strongest open coder. Requires very large multi-GPU deployment."},{"id":"llama-4-scout","name":"Llama 4 Scout","params_total":109,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":36,"swe_verified":68,"livecodebench":66,"aime":78,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-3xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-4xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":32,"tok_per_sec":22},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":45,"tok_per_sec":32},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"10M token context (longest in any open model). 109B / 17B active MoE."},{"id":"kimi-k2.6","name":"Kimi K2.6","params_total":235,"params_active":21,"released":"2026-05-01","license":"Modified MIT","jurisdiction":"China","swe_pro":47,"swe_verified":75,"livecodebench":78,"aime":84,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":36},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":32},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Best open-weight for sub-agent fan-out. Built for harness-driven parallel pipelines. Chinese-trained."},{"id":"glm-5.1","name":"GLM 5.1","params_total":358,"params_active":32,"released":"2026-04-22","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":76,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT license — rare among open frontier models besides DeepSeek. Strong for enterprise fine-tuning."},{"id":"mistral-small-4","name":"Mistral Small 4","params_total":24,"params_active":24,"released":"2026-04-18","license":"Apache 2.0","jurisdiction":"EU","swe_pro":28,"swe_verified":60,"livecodebench":58,"aime":68,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-3xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-4xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":48,"tok_per_sec":38},"mbp-16-m5max-128":{"quant":"BF16","vram_used":48,"tok_per_sec":55},"mba-15-32":{"quant":"Q4_K_M","vram_used":14,"tok_per_sec":30}},"notes":"6.5B effective parameters. EU jurisdiction. Best on-device option."},{"id":"nemotron-3-ultra-550b-a55b-moe","name":"Nvidia Nemotron 3 Ultra","params_total":550,"params_active":55,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q3_K_M","vram_used":180,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Hybrid Mamba-Transformer. 1M context. NVIDIA Open Model License. Computex June 4 launch."},{"id":"nemotron-3-nano-30b-a3b","name":"Nvidia Nemotron 3 Nano 30B-A3B","params_total":30,"params_active":3,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":30,"swe_verified":70,"livecodebench":68,"aime":76,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":24,"tok_per_sec":70},"mbp-16-m5max-128":{"quant":"BF16","vram_used":60,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":17,"tok_per_sec":50}},"notes":"Hybrid Mamba-Transformer Nano variant. Open weights."},{"id":"nemotron-nano-9b-v2","name":"Nvidia Nemotron Nano 9B v2","params_total":9,"params_active":9,"released":"2026-04-12","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":22,"swe_verified":60,"livecodebench":55,"aime":65,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-3xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-4xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":18,"tok_per_sec":90},"mbp-16-m5max-128":{"quant":"BF16","vram_used":18,"tok_per_sec":120},"mba-15-32":{"quant":"Q4_K_M","vram_used":6,"tok_per_sec":65}},"notes":"Dense 9B. Strong instruction following at edge sizes."},{"id":"kimi-k2.6-1t","name":"Kimi K2.6 (1T MoE)","params_total":1000,"params_active":32,"released":"2026-04-20","license":"Modified MIT","jurisdiction":"China","swe_pro":58.6,"swe_verified":80.2,"livecodebench":82,"aime":88,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Top open intelligence. 80.2% SWE-Verified, 58.6% SWE-Pro, AAII 54. 262K context. Modified MIT."},{"id":"glm-4.6","name":"GLM 4.6","params_total":355,"params_active":32,"released":"2025-09-15","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":75,"aime":80,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":28},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":24},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 200K context. Z.ai release."},{"id":"qwen3.6-35b-a3b","name":"Qwen 3.6 35B-A3B","params_total":35,"params_active":3,"released":"2026-04-16","license":"Apache 2.0","jurisdiction":"China","swe_pro":38,"swe_verified":74,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":22,"tok_per_sec":80},"mbp-16-m5max-128":{"quant":"BF16","vram_used":70,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"Apache 2.0 MoE — strong $/quality for fast bulk inference."},{"id":"cohere-command-a-plus","name":"Cohere Command A+","params_total":218,"params_active":28,"released":"2026-05-22","license":"CC-BY-NC + commercial","jurisdiction":"Canada","swe_pro":42,"swe_verified":75,"livecodebench":70,"aime":78,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":30},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":26},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"First open-weights Cohere release in &gt;1 year. Sparse-MoE multimodal."},{"id":"mistral-large-3-open","name":"Mistral Large 3 (open weights)","params_total":675,"params_active":41,"released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q3_K_M","vram_used":140,"tok_per_sec":22},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":180,"tok_per_sec":20},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"675B/41B active MoE. Apache 2.0. EU jurisdiction."},{"id":"deepseek-v4-flash-open","name":"DeepSeek V4-Flash (open weights)","params_total":284,"params_active":13,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":45},"vast-3xa6000":{"quant":"Q6_K","vram_used":115,"tok_per_sec":40},"vast-4xa6000":{"quant":"Q8_0","vram_used":150,"tok_per_sec":34},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":20},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 1M context, 384K max output. Cheap large-context open option."}],"frameworks":[{"name":"llama.cpp","best_for":"Mac (MLX), broad GGUF support","notes":"Best Apple Silicon performance via Metal."},{"name":"vLLM","best_for":"Multi-GPU servers, throughput","notes":"Production serving. Tensor parallelism."},{"name":"Ollama","best_for":"Easiest setup, dev workflow","notes":"Wrapper around llama.cpp. One-line model pull."},{"name":"MLX","best_for":"Apple Silicon native","notes":"Apple's framework. Best M-series perf."}]},"strategy":{"current_recommendation":{"label":"Risk-tiered GPT-5.6 + Fable stack","monthly_usd":null,"components":["GPT-5.6 Luna low — high-volume subagents and routine transformations","Gemini 3.5 Flash-Lite minimal — throughput-first extraction, parsing and parallel subagents","Gemini 3.6 Flash medium — Google-first coding, computer-use and multimodal agent loops","GPT-5.6 Terra medium — default engineering, review and documentation","GPT-5.6 Sol high or Claude Fable 5 high — escalation for hard, long-horizon tasks","Qwen 3.6 35B-A3B or DeepSeek V4-Flash — local route for privacy-sensitive bulk work"],"rationale":"Route by task risk instead of one subscription. Terra and Luna now retain strong coding-agent quality at substantially lower list price; Fable remains the long-horizon option only where its 30-day retention requirement is acceptable. Monthly cost is workload-dependent and must be computed from measured token volume."},"alternatives":[{"label":"Dual subscription (Claude Max + OpenAI Pro)","monthly_usd":230,"rationale":"Adds GPT-5.4 SWE-bench Pro lead and Codex CLI cloud sandbox. Worth $100/mo only if you frequently hit hard issues where Sonnet 4.6 plateaus.","verdict":"Defer until you have measured Sonnet plateau frequency for 30 days."},{"label":"API-only (no subscriptions)","monthly_usd":200,"rationale":"Pure pay-per-use. Maximum flexibility. Loses the subscription cache advantage — same workload costs 1.5–5× more for power users.","verdict":"Worse economics for your usage volume. Skip."},{"label":"Self-hosted maximalist","monthly_usd":80,"rationale":"Vast.ai 24/7 with Gemma 4 + Qwen 3 + occasional API top-up for frontier-only tasks.","verdict":"Cheapest if quality plateau at ~88% of Opus is acceptable. Operational overhead is real."}],"routing":[{"tier":"Bulk (70%)","use_for":"Classification, simple edits, triage, log parsing","preferred":"Gemini 3.5 Flash-Lite minimal · GPT-5.6 Luna low · Qwen 3.6 35B-A3B when local","cost_label":"Gemini $0.30/$2.50; Luna $0.20/$1.20 Mtok before cache, batch and reasoning tokens"},{"tier":"Mid (25%)","use_for":"Multi-file edits, code review, refactors, docs","preferred":"GPT-5.6 Terra medium · Gemini 3.6 Flash medium for Google-first or multimodal work","cost_label":"Terra $2/$12; Gemini $1.50/$7.50 Mtok before cache, batch and reasoning tokens"},{"tier":"Premium (5%)","use_for":"Architecture, hard debugging, long-context refactors","preferred":"GPT-5.6 Sol high · Claude Fable 5 high for long-horizon autonomy","cost_label":"Sol $5/$30; Fable $10/$50 Mtok before reasoning tokens"}],"open_questions":["Will the Nvidia Nemotron Coalition (Mistral, Cursor, Black Forest Labs, Thinking Machines) actually deliver a frontier-open consortium, or fragment within a quarter?","How does the June 15 Anthropic billing restructure (Chat pool + Agent SDK credit pool) change Max economics for agentic workloads?","Does Gemini 3.6 Flash’s lower output price and stronger agentic performance displace 3.1 Pro for most coding workflows before Gemini 3.5 Pro arrives?","DeepSeek V4-Pro permanent pricing ($0.435/$0.87) — does it force closed providers (OpenAI, Anthropic) to cut headline prices?","SubQ 1M-Preview's subquadratic architecture — does the next year see other commercial non-transformer LLMs?"],"task_fit":{"description":"For each task type, one recommended pick per provider drawn from the full roster. Click any cell to see runners-up that the daily briefing also considered. Burn × is derived from the matrix and recomputes when you change the reference dropdown.","rows":[{"task":"Tiny edits / boilerplate","description":"Single-file syntax fixes, regex replacements, formatting, docstrings, import sorting","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fixed-rate, fastest paid Anthropic option"},"openai":{"model_id":"gpt-5.5","effort":"minimal","rationale":"Cheapest reasoning setting on flagship"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Fastest Google route for cheap, high-volume pattern work"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Effectively $0/tok on existing Vast.ai infra"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ minimal","gpt-5.4-nano @ minimal"],"anthropic":["sonnet-4.6 @ low"],"google":["gemini-3.1-pro @ off"],"self_hosted":["llama-3.1-70b @ Q6_K"]}},{"task":"Normal coding task","description":"Single-function implementation, straightforward bug fixes, simple feature additions","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best $/quality on Anthropic side; cache reads included in Max 5×"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-specialised, ~⅓ burn of GPT-5.5 for near-identical SWE-Pro"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"58.7% SWE-Pro and stronger agentic coding than 3.5 Flash at lower output price"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Highest-quality open Gemma at full BF16 in 96GB"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium","gpt-5.5 @ low"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ high"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Multi-file implementation","description":"Feature spanning 3-10 files, requires understanding cross-file context","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Best style/intent understanding for code spanning files"},"openai":{"model_id":"gpt-5.5","effort":"medium","rationale":"Strong on planning; medium gives consistent multi-file edits"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Improved multi-step coding loops, fewer unwanted edits and 1M context"},"self_hosted":{"model_id":"qwen-3.6-plus","effort":null,"rationale":"Largest open MoE; strong reasoning across files"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ high","gpt-5.4 @ high"],"anthropic":["opus-4.7 @ high"],"google":["gemini-3.1-pro @ high"],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Hard debugging / architecture","description":"Race conditions, performance bottlenecks, design decisions with long-term implications","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Default Claude Code effort for Opus; best long-context retention"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"SWE-Pro leader; high effort balances depth and burn"},"google":{"model_id":"gemini-3.1-pro","effort":"high","rationale":"Preview Pro remains the deepest Google reasoning route; compare 3.6 Flash before paying the premium"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open-weight models trail frontier ~10-15pts here; not yet ready for hardest cases"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh","gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ high (lower burn)"],"google":[],"self_hosted":["qwen-3-235b-a22b — viable for some cases"]}},{"task":"Very hard autonomous repo task","description":"Overnight runs, agentic loops, tasks you don't supervise turn-by-turn","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"max","rationale":"Autonomy earns max effort's premium; no human in the loop to correct"},"openai":{"model_id":"gpt-5.5","effort":"xhigh","rationale":"Highest available reasoning budget on OpenAI side"},"google":{"model_id":null,"effort":null,"rationale":"Gemini 3.6 Flash improves agent loops but still lacks a max-equivalent tier for the hardest unsupervised runs"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Not recommended — frontier models earn their cost on the hardest tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ xhigh — better but $100/mo gating"],"anthropic":["opus-4.7 @ xhigh — if max budget is too steep"],"google":[],"self_hosted":[]}},{"task":"Code review / PR comments","description":"Reviewing diffs, suggesting improvements, finding issues without writing code yourself","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Critical thinking at moderate burn; doesn't need Opus depth"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-tuned reviewer at ⅓ burn of GPT-5.5"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Better instruction following and fewer unwanted code edits than 3.5 Flash"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Critique tasks don't need frontier; quality plateau acceptable"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ medium"],"self_hosted":["qwen-3-235b-a22b"]}},{"task":"Long-context refactor (50K+ tokens)","description":"Cross-cutting changes across a large codebase, where retention is the bottleneck","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"high","rationale":"1M context with retention that actually works; Opus's sweet spot"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"400K API context; degrades faster than Opus past 200K"},"google":{"model_id":"gemini-3.6-flash","effort":"high","rationale":"Google reports a 54% 1M-context MRCR score, roughly double 3.5 Flash"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open models cap around 200K useful context; not yet competitive"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh — if context fits in 256K"],"anthropic":["opus-4.7 @ xhigh"],"google":["gemini-3.1-pro @ high"],"self_hosted":[]}},{"task":"Subagent / delegated subtask","description":"Spawned by a main agent for focused work; quality threshold lower than user-facing","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fast, cheap, no effort knob to manage"},"openai":{"model_id":"gpt-5.4-mini","effort":"medium","rationale":"Best subagent — 94% of GPT-5.4 coding at 6× less"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Purpose-built for high-volume subagents at roughly 490 output tok/s"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Zero per-token cost for high-volume subagent loops"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ low","gpt-5.4-nano @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["llama-3.1-70b @ Q4_K_M"]}},{"task":"Documentation / explanation","description":"READMEs, code comments, tutorials, explaining existing code in prose","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best prose style on Anthropic side"},"openai":{"model_id":"gpt-5.4","effort":"medium","rationale":"Reasoning helps less here; medium-effort GPT-5.4 wins on $/quality"},"google":{"model_id":"gemini-3.6-flash","effort":"low","rationale":"Lower output price than 3.5 Flash with improved knowledge-work and document analysis"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Open models are competitive for prose tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ low"],"anthropic":["haiku-4.5"],"google":["gemini-3.5-flash-lite @ low"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Test scaffolding / fixtures","description":"Boilerplate test files, fixture generation, mocking, parameterised test suites","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Pattern-following work; doesn't need reasoning depth"},"openai":{"model_id":"gpt-5.4-mini","effort":"low","rationale":"Pattern-heavy; minimal reasoning, fast turnaround"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Lowest-latency current Google model for repetitive structured generation"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Bulk test scaffolding is exactly where self-hosted earns out"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Cybersecurity / vulnerability research","description":"Penetration testing, red-teaming, vulnerability discovery — tasks where cyber-specific safeguards apply","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Most capable Anthropic model; cyber tasks require Cyber Verification Program"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"Flagship reasoning; vetted security access via OpenAI"},"google":{"model_id":null,"effort":null,"rationale":"No cyber-specialized Gemini variant"},"self_hosted":{"model_id":"deepseek-v4-pro","effort":null,"rationale":"Open-weight without safety filtering; deploy in isolated environment"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ xhigh"],"google":[],"self_hosted":["qwen-3.6-plus","kimi-k2.6"]}}]}},"quota_burn_matrix":{"baseline_label":"selected reference model at medium effort = 1.00×","baseline_model_id":"fable-5","baseline_effort":"medium","methodology":"Burn ratios are displayed as multiples of the selected reference model at medium effort. Underlying values stored as absolute units anchored to gpt-5.5 medium = 1.00. Per-model ratios from API list pricing (firm, ±5%). Per-effort multipliers grounded in: Anthropic's published thinking-token budgets (low=skip, medium=~1k, high=5k, xhigh=10k, max=20k tokens); ArtificialAnalysis's measurement of Sonnet 4.6 max ≈ 3× Sonnet 4.5 on the Intelligence Index; nxcode.io / OpenAI guidance that xhigh ≈ 3-5× medium; ampcode's GPT-5.5 cost analysis. Per-effort multipliers are ±20%. NOTE: quality vs effort is not strictly monotonic per task — aggregate quality trends upward, but individual tasks can see high beat xhigh or medium beat high. OpenAI explicitly warns that 'high is not automatically better than medium'. Models that run in a separate preview quota bucket with no standard-pool burn are encoded as 0.00x and annotated until final token or credit rates are published.","stacking_multipliers":[{"name":"Fast mode (/fast on)","multiplier":2,"scope":"any cell"},{"name":"Cached input (long thread)","multiplier":0.6,"scope":"any cell"},{"name":"Plan mode (/plan)","multiplier":"uses high effort","scope":"overrides default"}],"openai_matrix":[{"model_id":"gpt-5.6-sol","minimal":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4,"note":"List-price blend equals the historical GPT-5.5 medium unit. Effort factors are planning estimates until measured token use is available."},{"model_id":"gpt-5.6-terra","minimal":0.18,"low":0.26,"medium":0.4,"high":0.68,"xhigh":1.08,"max":1.6,"note":"Two-fifths of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.6-luna","minimal":0.018,"low":0.026,"medium":0.04,"high":0.068,"xhigh":0.108,"max":0.16,"note":"Four percent of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.5","minimal":0.2,"low":0.5,"medium":1,"high":2,"xhigh":3.5},{"model_id":"gpt-5.5-pro","minimal":null,"low":null,"medium":6,"high":12,"xhigh":21},{"model_id":"gpt-5.4","minimal":0.1,"low":0.25,"medium":0.5,"high":1,"xhigh":1.75},{"model_id":"gpt-5.4-mini","minimal":0.011,"low":0.027,"medium":0.053,"high":0.107,"xhigh":0.187},{"model_id":"gpt-5.4-nano","minimal":0.003,"low":0.007,"medium":0.013,"high":null,"xhigh":null},{"model_id":"gpt-5.3-codex","minimal":0.07,"low":0.17,"medium":0.33,"high":0.67,"xhigh":1.17},{"model_id":"gpt-5.2","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null},{"model_id":"gpt-5.2-codex","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null}],"anthropic_matrix":[{"model_id":"fable-5","low":1.1,"medium":1.69,"high":2.87,"xhigh":4.56,"max":6.76,"note":"Always-on adaptive thinking. Values start from the $10/$50 list-price blend; effort multipliers are planning estimates rather than published token budgets."},{"model_id":"opus-4.8","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Current flagship. Same headline $5/$25. Fast Mode 3× cheaper ($10/$50) vs Opus 4.7."},{"model_id":"opus-4.7","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Legacy as of May 28. Thinking tokens: low=skip, medium=~1k, high=5k, xhigh=10k, max=20k. New tokenizer +35% tokens vs 4.6."},{"model_id":"sonnet-4.6","low":0.25,"medium":0.5,"high":1.25,"xhigh":null,"max":3.5,"note":"API default high. No xhigh tier. AA measured max ≈ 3× Sonnet 4.5 cost."},{"model_id":"haiku-4.5","low":null,"medium":0.17,"high":null,"xhigh":null,"max":null,"note":"No effort control — fixed-rate model."}],"google_matrix":[{"model_id":"gemini-3.1-pro","off":0.13,"low":0.2,"medium":0.3,"high":0.5},{"model_id":"gemini-3.5-flash","off":0.02,"low":0.03,"medium":0.05,"high":0.09},{"model_id":"gemini-3.6-flash","off":0.017,"low":0.025,"medium":0.042,"high":0.076},{"model_id":"gemini-3.5-flash-lite","off":0.005,"low":0.008,"medium":0.014,"high":0.024}],"non_reasoning_note":"Models without reasoning controls: Haiku 4.5 (Anthropic matrix above as fixed rate), GPT-5.5 Instant (ChatGPT default since May 5), Mistral Medium 3.5, MiniMax M2.7, Kimi K2.6, GLM 4.6, Qwen 3.6 variants, all Llama 4 variants, Grok 4.1 Fast, SubQ 1M-Preview — burn at fixed rate per model regardless of effort knob. DeepSeek V4-Flash-0731 now exposes thinking and non-thinking modes.","effort_quality_factors":{"minimal":0.65,"low":0.85,"medium":0.94,"high":0.98,"xhigh":1,"max":1.02,"off":0.85,"on":0.95},"quality_methodology":"Quality % = (model_swe_pro / reference_swe_pro × effort_quality_factor) × 100. Effort quality factors: minimal=0.65, low=0.85, medium=0.94, high=0.98, xhigh=1.00, max=1.02. Anchors: Anthropic's Hex measurement (low Opus 4.7 ≈ medium Opus 4.6 quality) anchors low ≈ 0.85; apiyi.com's report that max gains ~3pts over xhigh on hardest tasks anchors max=1.02; ampcode's GPT-5.5 analysis (medium captures 'most' of capability) anchors medium=0.94. IMPORTANT: these are AGGREGATE estimates. stet.sh found per-task reversals — high can beat xhigh on some tasks, medium can beat high. Treat as ±10% indicators, not precise measures.","unit_anchor":"gpt-5.5 medium","unit_anchor_note":"All burn values are stored as absolute units anchored to gpt-5.5 medium = 1.00. At render time, each cell is divided by the reference model's medium-effort cell (or its single datapoint for non-reasoning models) to produce the displayed ×-of-reference number. When the reference dropdown changes, all burn ratios recalculate against the new reference.","workload_presets":{"cold":{"label":"Cold (0% cache)","description":"One-off prompts, no context reuse. Worst-case burn.","cache_hit_rate":0},"mixed":{"label":"Mixed (40% cache)","description":"Interactive coding with some context reuse. Typical default.","cache_hit_rate":0.4},"warm":{"label":"Warm (70% cache)","description":"Sustained agentic session. AA's industry-standard 7:2:1 blend assumes this rate. danielvaughan.com's Codex CLI model uses 70%.","cache_hit_rate":0.7},"hot":{"label":"Hot (90% cache)","description":"Long warmer-pattern sessions with stable system prompts and reused context. vsits.co documented 99% achievable on Claude Code Max subscription with optimization.","cache_hit_rate":0.9}},"default_workload":"mixed","cache_methodology":"Cache discount = (cache_hit_price / input_price). Lower = better discount. Effective burn = (1 - cache_hit_rate) × raw_burn + cache_hit_rate × raw_burn × cache_discount. Per-provider cache discounts: Anthropic 10% (90% off), OpenAI 25% on flagship/10% on 5.4 family, DeepSeek V4-Pro 0.83% (99.2% off — most aggressive in industry), DeepSeek V4-Flash 2%, MiniMax 52%, Mistral &amp; Grok no published cache pricing (modeled as 100% = no discount), Google ~25% plus per-hour storage fee (not modeled). IMPORTANT CAVEAT on Anthropic subscriptions: docs say cache reads count at 10% rate, but GitHub issue anthropics/claude-code#24147 reports cache reads burning quota at full rate on Max subscriptions. Treat the displayed cache benefit as accurate for API usage; subscription quota behavior is contested. Cache writes (1.25× / 2× of input rate on Anthropic) are not modeled — amortize across many reads in steady-state."},"sources":[{"n":1,"category":"Anthropic","title":"Anthropic news","url":"https://www.anthropic.com/news"},{"n":2,"category":"Anthropic","title":"Anthropic Claude model overview","url":"https://docs.claude.com/en/docs/about-claude/models/overview"},{"n":3,"category":"Anthropic","title":"Anthropic pricing","url":"https://www.anthropic.com/pricing"},{"n":4,"category":"OpenAI","title":"OpenAI news","url":"https://openai.com/news/"},{"n":5,"category":"OpenAI","title":"OpenAI model catalog","url":"https://platform.openai.com/docs/models"},{"n":6,"category":"OpenAI","title":"OpenAI API pricing","url":"https://openai.com/api/pricing/"},{"n":7,"category":"Google","title":"Google DeepMind blog","url":"https://blog.google/technology/google-deepmind/"},{"n":8,"category":"Google","title":"Gemini API models","url":"https://ai.google.dev/gemini-api/docs/models"},{"n":9,"category":"Google","title":"Gemini API pricing","url":"https://ai.google.dev/pricing"},{"n":10,"category":"Mistral","title":"Mistral news","url":"https://mistral.ai/news/"},{"n":11,"category":"Mistral","title":"Mistral La Plateforme pricing","url":"https://mistral.ai/products/la-plateforme#pricing"},{"n":12,"category":"xAI","title":"xAI blog (Grok)","url":"https://x.ai/blog"},{"n":13,"category":"DeepSeek","title":"DeepSeek API docs","url":"https://api-docs.deepseek.com/"},{"n":14,"category":"MiniMax","title":"MiniMax news","url":"https://www.minimaxi.com/en/news"},{"n":15,"category":"Benchmarks","title":"SWE-bench (Verified)","url":"https://www.swebench.com/","note":"SWE-bench Verified is widely considered contaminated; treat with skepticism."},{"n":16,"category":"Benchmarks","title":"SWE-bench Pro leaderboard","url":"https://scale.com/leaderboard/swe_bench_pro","note":"Preferred trustworthy code-agent benchmark."},{"n":17,"category":"Benchmarks","title":"LiveCodeBench","url":"https://livecodebench.github.io/"},{"n":18,"category":"Benchmarks","title":"Artificial Analysis (cross-provider pricing + benchmarks)","url":"https://artificialanalysis.ai/"},{"n":19,"category":"Benchmarks","title":"LMArena (chatbot arena)","url":"https://lmarena.ai/"},{"n":20,"category":"Open-weight","title":"Hugging Face — trending models","url":"https://huggingface.co/models?sort=trending"},{"n":21,"category":"Open-weight","title":"Hugging Face blog","url":"https://huggingface.co/blog"},{"n":22,"category":"Open-weight","title":"r/LocalLLaMA (community signal)","url":"https://www.reddit.com/r/LocalLLaMA/"},{"n":23,"category":"Harnesses","title":"Claude Code (anthropics/claude-code)","url":"https://github.com/anthropics/claude-code"},{"n":24,"category":"Harnesses","title":"OpenCode (sst/opencode)","url":"https://github.com/sst/opencode"},{"n":25,"category":"Harnesses","title":"Codex CLI (openai/codex)","url":"https://github.com/openai/codex"},{"n":26,"category":"Harnesses","title":"Gemini CLI (google-gemini/gemini-cli)","url":"https://github.com/google-gemini/gemini-cli"},{"n":27,"category":"Harnesses","title":"Aider (Aider-AI/aider)","url":"https://github.com/Aider-AI/aider"},{"n":28,"category":"Harnesses","title":"Cline (cline/cline)","url":"https://github.com/cline/cline"},{"n":29,"category":"Harnesses","title":"Roo Code (RooVetGit/Roo-Code)","url":"https://github.com/RooVetGit/Roo-Code"},{"n":30,"category":"Harnesses","title":"Goose (block/goose)","url":"https://github.com/block/goose"},{"n":31,"category":"Harnesses","title":"Cursor changelog","url":"https://cursor.com/changelog"},{"n":32,"category":"Harnesses","title":"Windsurf changelog","url":"https://windsurf.com/changelog"},{"n":33,"category":"Hardware","title":"Vast.ai (spot GPU pricing — A6000, H100)","url":"https://vast.ai/"},{"n":34,"category":"Hardware","title":"Apple MacBook Pro (M-series)","url":"https://www.apple.com/shop/buy-mac/macbook-pro"},{"n":35,"category":"Policy","title":"Anthropic OAuth third-party restrictions (Apr 4, 2026)","url":"https://www.anthropic.com/news"},{"n":36,"category":"Policy","title":"OpenAI Codex token-based pricing migration (Apr 2, 2026)","url":"https://openai.com/news/"},{"n":37,"category":"Aggregators","title":"Hacker News (AI tags)","url":"https://news.ycombinator.com/"},{"n":38,"category":"Benchmarks","title":"AIME (math competition benchmark)","url":"https://aimeproblems.com/","note":"Used to anchor Reasoning axis ratings."},{"n":39,"category":"Benchmarks","title":"MMLU / MMLU-Pro (knowledge benchmark)","url":"https://github.com/hendrycks/test","note":"Used to anchor Knowledge axis ratings."},{"n":40,"category":"Benchmarks","title":"tau-bench (tool-use + agentic behaviour)","url":"https://github.com/sierra-research/tau-bench","note":"Used to anchor Agentic axis ratings."},{"n":41,"category":"Nvidia","title":"build.nvidia.com","url":"https://build.nvidia.com/"},{"n":42,"category":"Nvidia","title":"Hugging Face — Nvidia","url":"https://huggingface.co/nvidia"},{"n":43,"category":"Nvidia","title":"Nvidia blogs","url":"https://blogs.nvidia.com/"},{"n":44,"category":"Benchmarks","title":"SWE-bench Pro public leaderboard (Scale)","url":"https://scale.com/leaderboard/swe_bench_pro_public"},{"n":45,"category":"Benchmarks","title":"Scale Labs","url":"https://labs.scale.com/"},{"n":46,"category":"Benchmarks","title":"SWE-Rebench","url":"https://swe-rebench.com/"},{"n":47,"category":"Benchmarks","title":"Terminal-Bench 2.0","url":"https://tbench.ai/leaderboard"},{"n":48,"category":"Aggregators","title":"OpenRouter","url":"https://openrouter.ai/"},{"n":49,"category":"Aggregators","title":"LLM-Stats","url":"https://llm-stats.com/"},{"n":50,"category":"Benchmarks","title":"Artificial Analysis Intelligence Index","url":"https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index"},{"n":51,"category":"Harnesses","title":"Claude Code changelog","url":"https://code.claude.com/docs/en/changelog"},{"n":52,"category":"Harnesses","title":"Cursor pricing","url":"https://cursor.com/pricing"},{"n":53,"category":"Harnesses","title":"Windsurf changelog (Devin Desktop)","url":"https://windsurf.com/changelog"},{"n":54,"category":"Harnesses","title":"Gemini CLI changelogs","url":"https://geminicli.com/docs/changelogs"},{"n":55,"category":"Harnesses","title":"Codex CLI releases","url":"https://github.com/openai/codex/releases"},{"n":56,"category":"Google","title":"DeepMind models","url":"https://deepmind.google/models"},{"n":57,"category":"Mistral","title":"Mistral pricing","url":"https://mistral.ai/pricing"},{"n":58,"category":"Google","title":"Gemini API pricing (canonical)","url":"https://ai.google.dev/gemini-api/docs/pricing"},{"n":59,"category":"MiniMax","title":"Artificial Analysis — MiniMax M2.7","url":"https://artificialanalysis.ai/models/minimax-m2-7"},{"n":60,"category":"Moonshot","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3"},{"n":61,"category":"Moonshot","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models"},{"n":62,"category":"OpenAI","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/"},{"n":63,"category":"Hosting","title":"Runpod GPU cloud pricing","url":"https://www.runpod.io/pricing"},{"n":64,"category":"Hosting","title":"Runpod July 2024 GPU price changes","url":"https://www.runpod.io/blog/runpod-slashes-gpu-prices-more-power-less-cost-for-ai-builders"},{"n":65,"category":"Hosting","title":"Contabo GPU Cloud configurations and pricing","url":"https://contabo.com/en/gpu-cloud/"},{"n":66,"category":"Hosting","title":"Infomaniak Public Cloud pricing","url":"https://www.infomaniak.com/en/hosting/public-cloud/prices"},{"n":67,"category":"Hosting","title":"Infomaniak GPU flavor documentation","url":"https://docs.infomaniak.cloud/compute/instances/flavors/"},{"n":68,"category":"Hosting","title":"Hyperstack GPU cloud pricing","url":"https://www.hyperstack.cloud/"},{"n":69,"category":"Hosting","title":"Vast.ai marketplace pricing methodology","url":"https://docs.vast.ai/guides/instances/pricing"},{"n":70,"category":"Harnesses","title":"Codex CLI 0.144.6 release","url":"https://github.com/openai/codex/releases/tag/rust-v0.144.6"},{"n":71,"category":"Harnesses","title":"Gemini CLI 0.51.0 release","url":"https://github.com/google-gemini/gemini-cli/releases/tag/v0.51.0"},{"n":72,"category":"Harnesses","title":"OpenCode 1.18.4 release","url":"https://github.com/anomalyco/opencode/releases/tag/v1.18.4"},{"n":73,"category":"Harnesses","title":"Cline 4.0.10 release","url":"https://github.com/cline/cline/releases/tag/v4.0.10"},{"n":74,"category":"Harnesses","title":"Goose 1.43.0 release","url":"https://github.com/aaif-goose/goose/releases/tag/v1.43.0"},{"n":75,"category":"Moonshot","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing"},{"n":76,"category":"Swiss AI Initiative","title":"Apertus v1.1 4B Instruct model card","url":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"n":77,"category":"Google","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"n":78,"category":"Google","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"n":79,"category":"Google","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"n":80,"category":"Google","title":"Gemini API release notes","url":"https://ai.google.dev/gemini-api/docs/changelog"},{"n":81,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash"},{"n":82,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"},{"n":83,"category":"OpenAI","title":"GPT-5.6 launch and benchmark table","url":"https://openai.com/index/gpt-5-6/"},{"n":84,"category":"Google","title":"Gemini API Managed Agents update","url":"https://blog.google/innovation-and-ai/technology/developers-tools/expanding-managed-agents-gemini-api/"},{"n":85,"category":"Benchmarks","title":"Scale Labs SWE-bench Pro public leaderboard","url":"https://labs.scale.com/api/pdf/leaderboard/swe_bench_pro_public","note":"Public leaderboard values use their own model versions and evaluation setup; do not merge them mechanically with vendor launch tables."},{"n":86,"category":"Anthropic","title":"What's new in Claude Opus 5","url":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5","note":"Official launch specifications, pricing, availability, and migration behaviour."},{"n":87,"category":"Benchmarks","title":"Artificial Analysis — Claude Opus 5","url":"https://artificialanalysis.ai/models/claude-opus-5","note":"Independent effort-specific intelligence, latency, throughput, and price analysis."},{"n":88,"category":"Thinking Machines Lab","title":"Inkling model card","url":"https://thinkingmachines.ai/model-card/inkling/"},{"n":89,"category":"OpenAI","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra"},{"n":90,"category":"OpenAI","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna"},{"n":91,"category":"DeepSeek","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","note":"Vendor-reported benchmark results use DeepSeek Harness minimal mode at max effort; DSBench results are internal."},{"n":92,"category":"DeepSeek","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/"},{"n":93,"category":"OpenAI","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","note":"Official July 30 pricing, paid-subscription credit, and Sol API Fast mode details."},{"n":94,"category":"OpenAI","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","note":"Official August 6 ChatGPT product update; it does not change the API model contract."}],"provider_sources":{"Anthropic":[1,2,3,86,87],"OpenAI":[4,5,6,62,83,89,90,93,94],"Google":[7,8,9,58,77,78,79,80,81,82,84],"Mistral":[10,11],"xAI":[12],"DeepSeek":[13,91,92],"MiniMax":[14],"Meta":[20,21],"Alibaba":[20,21],"Moonshot":[20,21,60,61,75],"Zhipu":[20,21],"Nvidia":[41,42,43],"Cohere":[20,21],"Swiss AI Initiative":[76],"SubQ":[],"Thinking Machines Lab":[88]},"section_sources":{"models":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,60,61,62,76,77,78,79,80,81,82,86,87,89,90,91,92],"harnesses":[23,24,25,26,27,28,29,30,31,32,35,51,52,53,54,55,70,71,72,73,74],"self_hosting":[20,21,22,33,34,63,64,65,66,67,68,69,76],"strategy":[1,3,4,6,16,18,33,35,36]},"capabilities":{"schema_version":"1.0","axes":[{"key":"coding","label":"Coding","short":"Code","effort_sensitivity":0.5},{"key":"reasoning","label":"Reasoning &amp; Architecture","short":"R&amp;A","effort_sensitivity":0.6},{"key":"knowledge","label":"Knowledge &amp; Research","short":"K&amp;R","effort_sensitivity":0.05},{"key":"comms","label":"Communication &amp; Docs","short":"Comms","effort_sensitivity":0.2},{"key":"multimodal","label":"Multimodal","short":"MM","effort_sensitivity":0.05},{"key":"agentic","label":"Agentic","short":"Agent","effort_sensitivity":0.4}],"level_labels":{"coding":{"1":"Snippet","2":"Standard","3":"Cross-file","4":"Hard","5":"Frontier"},"reasoning":{"1":"Apply known","2":"Multi-step","3":"Cross-cutting","4":"Novel system","5":"Research-grade"},"knowledge":{"1":"Recall","2":"Contextual","3":"Cross-domain","4":"Frontier","5":"Original"},"comms":{"1":"Grammatical","2":"Structured","3":"Tutorial","4":"Editorial","5":"Publishable"},"multimodal":{"1":"Text only","2":"Image-in","3":"Image reasoning","4":"Video/audio","5":"Cross-modal gen"},"agentic":{"1":"Single-turn","2":"Tool calls","3":"Plan coherence","4":"Self-correcting","5":"Long-horizon"}},"focus_presets":{"balanced":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"coding-focused":{"coding":2,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"architecture-focused":{"coding":1,"reasoning":2,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"research-focused":{"coding":1,"reasoning":1,"knowledge":2,"comms":1,"multimodal":1,"agentic":1},"writing-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":2,"multimodal":1,"agentic":1},"multimodal-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":2,"agentic":1},"agentic-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":2}},"default_focus":"balanced","methodology":"1-5 levels per axis; Coding rated against SWE-Pro / SWE-Verified evidence; Reasoning against AIME / LiveCodeBench / public hard-reasoning evals; Multimodal against published modality support (image/audio/video in/out); Agentic rated holding harness constant at Claude-Code-class baseline (plan coherence, tool-call quality, self-correction, calibrated stopping, refusal hygiene); Knowledge from model-card claims + community-validated state-of-art awareness; Comms from prose-evaluation rounds + structural quality. Ratings carry citation; see src/dashboard-context.md for the full rubric.","effort_formula":"effective(axis, effort) = max(0, ceiling[axis] × (1 − effort_sensitivity[axis] × (1 − radar_effort_factor[effort]))). radar_effort_factor[medium] = 1.00 → primary polygon equals reference polygon shape at medium effort for the same model. Lower efforts shrink the polygon by sensitivity-weighted amounts; higher efforts grow it and may push vertices OUTSIDE the outer level-50 ring — the ring is a rubric marker, not a hard cap. Vertices that exceed 50 are drawn in accent-hot to signal 'boosted above medium-effort baseline'.","radar_effort_factors":{"minimal":0.5,"low":0.75,"medium":1,"high":1.08,"xhigh":1.15,"max":1.22},"scale":{"rubric_min":1,"rubric_max":50,"visual_max":60,"band_size":10,"boost_band":"51–60","interpretation":"capability_levels[axis] = the model's effective capability at MEDIUM effort. Briefings rate from 1 to 50. The radar's visual scale extends to 60 to accommodate the 'effort-boost band' (51–60) — values that fall there are formula-derived from extra reasoning effort, never directly rated. The outer rubric ring at 50 is visually emphasized as 'current frontier'.","bands":{"1-10":"Snippet","11-20":"Standard","21-30":"Cross-file","31-40":"Hard","41-50":"Frontier","51-60":"Effort-boost band (derived, not rated)"}}},"report_metrics":{"schema_version":"1.0","researched_at":"2026-08-01","currency":"USD","reference_options":["fable-5","gpt-5.6-sol","gpt-5.6-terra","gpt-5.6-luna","opus-4.8","gpt-5.5","gpt-5.5-pro","gpt-5.5-instant","kimi-k3","gemini-3.6-flash","gemini-3.5-flash-lite"],"default_reference":"fable-5","methodology":{"benchmark_scale":"Each benchmark is divided by benchmark.max_value and expressed on a 0-100 scale. The quality composite is the weighted arithmetic mean of available normalized benchmarks; it is shown only when coverage is at least 0.50.","quality_weights":{"aaii_v4_1":0.2,"coding_agent_index_v1_1":0.2,"swe_bench_pro":0.2,"deep_swe_v1_1":0.15,"terminal_bench_2_1":0.15,"agents_last_exam":0.1},"speed_score":"100 * sqrt(task_speed_index / max_task_speed_index). task_speed_index is 100 * reference_time / model_time. It is not output tokens per second and is comparable only within the cited evaluation setup.","cost_score":"10 + 90 * ln(max_blended_price / blended_price) / ln(max_blended_price / min_blended_price). Blended price is 0.30 * input_price + 0.70 * output_price per million tokens. Higher cost score is better.","scq_compound":"Geometric mean of quality_score, speed_score and cost_score. Geometric mean prevents one category from fully compensating for a weak category. Null when quality coverage is below 0.50.","capability_compound":"Weighted arithmetic mean of the six capability axes. Balanced uses weight 1 for each axis; a selected focus uses weight 2.5 for that axis and 1 for every other axis.","reference_quality":"Each model stores or derives a Fable-anchored quality value. Displayed quality = 100 * model_anchor / selected_reference_anchor, so the selected reference is always exactly 100. New-suite composites use only explicitly overlapping evaluations; legacy rows fall back to SWE-Bench Pro. Unknown evidence remains null, never zero.","burn":"Raw blended-price units are multiplied by an effort factor and cache factor, then divided by the selected reference model at medium effort. Cache factor = (1-hit_rate) + hit_rate*cache_read_ratio.","missing_data":"Keep unknown values null. Display insufficient comparable evidence rather than zero. Detailed benchmark cells remain version-specific and may be empty even when a reference-relative quality anchor exists from a separate documented comparison set."},"market_signals":[{"id":"frontier_quality","label":"Frontier quality index","explanation":"Best broad intelligence score published on Artificial Analysis Intelligence Index v4.1. Higher is better; the benchmark version must remain fixed across the series.","unit":"index","direction":"up","current":59.9,"history":[{"date":"2026-01-01","value":48.2},{"date":"2026-02-01","value":50.1},{"date":"2026-03-01","value":52.7},{"date":"2026-04-01","value":55.7},{"date":"2026-05-01","value":55.7},{"date":"2026-06-01","value":59.9},{"date":"2026-07-01","value":59.9}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"coding_quality","label":"Coding-agent frontier","explanation":"Best score on Artificial Analysis Coding Agent Index v1.1. Higher means better end-to-end coding-agent performance, not just code completion.","unit":"index","direction":"up","current":80,"history":[{"date":"2026-01-01","value":66.1},{"date":"2026-02-01","value":69.8},{"date":"2026-03-01","value":72.5},{"date":"2026-04-01","value":76.4},{"date":"2026-05-01","value":77.2},{"date":"2026-06-01","value":77.2},{"date":"2026-07-01","value":80}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"task_speed","label":"Coding task speed increase","explanation":"Fastest current end-to-end coding task-rate index, where Fable 5 is fixed at 100. A value of 320 means an estimated 3.2 times as many comparable tasks per unit time; it is not tokens per second.","unit":"Fable=100","direction":"up","current":320,"history":[{"date":"2026-01-01","value":100},{"date":"2026-02-01","value":112},{"date":"2026-03-01","value":126},{"date":"2026-04-01","value":145},{"date":"2026-05-01","value":180},{"date":"2026-06-01","value":180},{"date":"2026-07-01","value":320}],"source":"https://openai.com/index/gpt-5-6/","history_note":"Pre-July points are frozen planning estimates from prior report snapshots and should not be interpreted as one controlled longitudinal benchmark."},{"id":"blended_frontier_price","label":"Lowest frontier blended API price","explanation":"Lowest 30% input / 70% output list-price blend among models meeting the report's frontier quality floor. Lower is better; batch and caching are excluded.","unit":"$/Mtok","direction":"down","current":0.9,"history":[{"date":"2026-01-01","value":18},{"date":"2026-02-01","value":16.5},{"date":"2026-03-01","value":14},{"date":"2026-04-01","value":11.4},{"date":"2026-05-01","value":9},{"date":"2026-06-01","value":9},{"date":"2026-07-01","value":4.5},{"date":"2026-08-01","value":0.9}],"source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"open_weight_quality","label":"Open-weight quality vs selected reference","explanation":"Best open-weight model relative to the selected reference. Stored history is Fable-anchored and rebased in the browser whenever the reference changes. July uses Kimi K3's geometric mean across 14 overlapping launch-suite values; vendor-reported values require independent reproduction.","unit":"%","direction":"up","current":99.2,"history":[{"date":"2026-01-01","value":82},{"date":"2026-02-01","value":85},{"date":"2026-03-01","value":88},{"date":"2026-04-01","value":91},{"date":"2026-05-01","value":94},{"date":"2026-06-01","value":96},{"date":"2026-07-01","value":99.2}],"history_note":"Earlier points are frozen report estimates; July uses Moonshot’s K3 launch suite and is not a single controlled longitudinal benchmark.","source":"https://www.kimi.com/blog/kimi-k3"},{"id":"output_throughput","label":"Documented API output throughput","explanation":"Fastest provider-documented general API output rate in the tracked set. Unit is generated tokens per second; it is distinct from time to first token and end-to-end task speed.","unit":"tok/s","direction":"up","current":490,"history":[{"date":"2026-01-01","value":110},{"date":"2026-02-01","value":130},{"date":"2026-03-01","value":140},{"date":"2026-04-01","value":180},{"date":"2026-05-01","value":200},{"date":"2026-06-01","value":220},{"date":"2026-07-01","value":490}],"history_note":"Provider and independent figures use different serving and context conditions. The latest point is Artificial Analysis's Gemini 3.5 Flash-Lite measurement on Google's first-party API.","source":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"}],"benchmarks":[{"id":"aaii_v4_1","label":"AA Intelligence Index v4.1","max_value":59.9,"unit":"index","good_direction":"high"},{"id":"coding_agent_index_v1_1","label":"AA Coding Agent Index v1.1","max_value":80,"unit":"index","good_direction":"high"},{"id":"swe_bench_pro","label":"SWE-Bench Pro","max_value":80,"unit":"%","good_direction":"high"},{"id":"deep_swe_v1_1","label":"DeepSWE v1.1","max_value":72.7,"unit":"%","good_direction":"high"},{"id":"terminal_bench_2_1","label":"Terminal-Bench 2.1","max_value":91.9,"unit":"%","good_direction":"high","note":"Maximum is GPT-5.6 Sol Ultra; ordinary Sol scores 88.8."},{"id":"agents_last_exam","label":"Agents' Last Exam","max_value":52.7,"unit":"%","good_direction":"high"}],"model_metrics":[{"model_id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","benchmarks":{"aaii_v4_1":59.9,"coding_agent_index_v1_1":77.2,"swe_bench_pro":80,"deep_swe_v1_1":69.7,"terminal_bench_2_1":83.1,"agents_last_exam":40.5},"quality_score":94.9,"quality_coverage":1,"task_speed_index":100,"speed_score":50,"api_input":10,"api_cached_input":1,"api_output":50,"blended_price":38,"cost_score":10,"scq_compound":36.2,"speed_evidence":"Reference index; Anthropic labels comparative latency slower.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","benchmarks":{"aaii_v4_1":58.9,"coding_agent_index_v1_1":80,"swe_bench_pro":64.6,"deep_swe_v1_1":72.7,"terminal_bench_2_1":88.8,"agents_last_exam":52.7},"quality_score":95.4,"quality_coverage":1,"task_speed_index":256,"speed_score":80,"api_input":5,"api_cached_input":0.5,"api_output":30,"blended_price":22.5,"cost_score":22.6,"scq_compound":55.7,"speed_evidence":"OpenAI introduced API Fast mode on July 30, claiming up to 2.5× Standard speed at 2× price with no intelligence change; the base task-speed index remains a separate vendor comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","benchmarks":{"aaii_v4_1":55,"coding_agent_index_v1_1":77.4,"swe_bench_pro":63.4,"deep_swe_v1_1":69.6,"terminal_bench_2_1":87.4,"agents_last_exam":50.4},"quality_score":91.8,"quality_coverage":1,"task_speed_index":300,"speed_score":86.6,"api_input":2,"api_cached_input":0.2,"api_output":12,"blended_price":9,"cost_score":44.6,"scq_compound":70.8,"speed_evidence":"OpenAI reports roughly one-third of Fable 5 task time for the family coding comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","benchmarks":{"aaii_v4_1":51.2,"coding_agent_index_v1_1":74.6,"swe_bench_pro":62.7,"deep_swe_v1_1":67.2,"terminal_bench_2_1":84.7,"agents_last_exam":50.3},"quality_score":88.7,"quality_coverage":1,"task_speed_index":320,"speed_score":89.4,"api_input":0.2,"api_cached_input":0.02,"api_output":1.2,"blended_price":0.9,"cost_score":100,"scq_compound":92.6,"speed_evidence":"OpenAI calls Luna the fastest tier; 320 is a conservative planning index above the cited 300 family comparison and must not be treated as measured tok/s.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","benchmarks":{"deep_swe_v1_1":67.5,"terminal_bench_2_1":88.3},"quality_score":null,"quality_coverage":0.33,"quality_vs_fable":99.2,"task_speed_index":null,"speed_score":null,"api_input":3,"api_cached_input":0.3,"api_output":15,"blended_price":11.4,"cost_score":38.9,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time or output-throughput figure published for K3 at launch.","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"model_id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","benchmarks":{"aaii_v4_1":55.7,"coding_agent_index_v1_1":72.5,"swe_bench_pro":69.2,"deep_swe_v1_1":59,"terminal_bench_2_1":78.9,"agents_last_exam":45.2},"quality_score":87.7,"quality_coverage":1,"task_speed_index":180,"speed_score":67.1,"api_input":5,"api_cached_input":0.5,"api_output":25,"blended_price":19,"cost_score":26.7,"scq_compound":54,"speed_evidence":"Planning index based on moderate vendor latency and optional fast mode; validate in the local harness.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gemini-3.1-pro","name":"Gemini 3.1 Pro Preview","provider":"Google","benchmarks":{"aaii_v4_1":46.5,"coding_agent_index_v1_1":42.7,"swe_bench_pro":54.2,"deep_swe_v1_1":11.8,"terminal_bench_2_1":70.7,"agents_last_exam":32.1},"quality_score":59.8,"quality_coverage":1,"task_speed_index":null,"speed_score":null,"api_input":null,"api_cached_input":null,"api_output":null,"blended_price":null,"cost_score":null,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time measurement in the selected source set.","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro"},{"model_id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","benchmarks":{"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78},"quality_score":77.4,"quality_coverage":0.7,"quality_vs_fable":81.6,"task_speed_index":null,"speed_score":null,"api_input":1.5,"api_cached_input":0.15,"api_output":7.5,"blended_price":5.7,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 304 output tok/s at high thinking; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"model_id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","benchmarks":{"aaii_v4_1":36,"swe_bench_pro":54.2,"terminal_bench_2_1":54},"quality_score":62.5,"quality_coverage":0.55,"quality_vs_fable":65.9,"task_speed_index":null,"speed_score":null,"api_input":0.3,"api_cached_input":0.03,"api_output":2.5,"blended_price":1.84,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 490 output tok/s; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"}],"visualizations":{"bubble":{"x":"cost_score","y":"speed_score","size":"quality_vs_selected_reference","color":"provider","include_when":"quality vs reference, speed score and cost score are all non-null"},"heatmap":{"rows":"model","columns":["quality_vs_selected_reference","speed_score","cost_score"],"color_scale":"sequential_good","domain":[0,100]},"bcg":{"x":"capability_compound","y":"mean(quality_vs_selected_reference, speed_score, cost_score)","color":"provider","quadrants":"medians of currently visible models"}},"capability":{"axes":["coding","reasoning_architecture","knowledge_research","communication_docs","multimodal","agentic"],"focus_multiplier":2.5,"models":[{"model_id":"fable-5","scores":[98,99,96,96,90,99],"capability_compound":96.3,"scq_compound":36.2},{"model_id":"gpt-5.6-sol","scores":[98,97,95,95,96,98],"capability_compound":96.5,"scq_compound":55.7},{"model_id":"gpt-5.6-terra","scores":[95,92,91,92,90,95],"capability_compound":92.5,"scq_compound":70.8},{"model_id":"gpt-5.6-luna","scores":[92,88,86,89,86,91],"capability_compound":88.7,"scq_compound":92.6},{"model_id":"opus-4.8","scores":[92,94,94,95,82,94],"capability_compound":91.8,"scq_compound":54},{"model_id":"gemini-3.1-pro","scores":[78,84,94,84,98,72],"capability_compound":85,"scq_compound":null}],"provenance_note":"Capability scores are editorial rubric ratings synthesized from benchmark and feature evidence, not vendor benchmark results. Keep them separate from benchmark values."},"economics":{"effort_factors":{"none":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4},"cache_hit_presets":{"cold":0,"mixed":0.4,"warm":0.7,"hot":0.9},"models":[{"model_id":"fable-5","base_blended_price":38,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"],"note":"Always-on adaptive thinking; effort is a behavioral control, not a published token multiplier."},{"model_id":"gpt-5.6-sol","base_blended_price":22.5,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-terra","base_blended_price":9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-luna","base_blended_price":0.9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"opus-4.8","base_blended_price":19,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"]},{"model_id":"kimi-k3","base_blended_price":11.4,"cache_read_ratio":0.1,"efforts":["max"],"note":"Always-on reasoning at launch; the max-only effort control is behavioral and not a published token multiplier."}],"efficiency_filters":{"quality_floor_default":75,"provider_default":"all","effort_default":"medium","workload_default":"mixed","sort_default":"scq_compound_desc"}},"hardware_options":[{"id":"local-64gb","name":"64 GB unified-memory workstation","memory_gb":64,"type":"local","best_for":"3B-35B dense or small MoE models at Q4-Q8","cost_note":"Capex varies; compare measured memory bandwidth, not product year."},{"id":"local-128gb","name":"128 GB unified-memory workstation","memory_gb":128,"type":"local","best_for":"Up to roughly 100 GB quantized weights with headroom for KV cache","cost_note":"Portable and quiet; slower than datacenter GPUs for sustained batches."},{"id":"cloud-2xa6000","name":"2 x RTX A6000 48 GB","memory_gb":96,"type":"cloud","best_for":"70B-class dense and mid-size MoE Q4 deployments","cost_note":"Spot price varies by host; record price and interconnect at test time."},{"id":"cloud-h200","name":"1 x H200 141 GB","memory_gb":141,"type":"cloud","best_for":"High-throughput 70B inference and larger quantized MoE models","cost_note":"Use provider quote; hourly rates change frequently."},{"id":"cloud-b200","name":"1 x B200 192 GB","memory_gb":192,"type":"cloud","best_for":"Large-model throughput where software stack supports Blackwell","cost_note":"Availability and hourly rates vary; validate framework support."}],"hardware_model_fit":[{"model_id":"qwen3.6-35b-a3b","hardware_id":"local-64gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":55,"evidence":"planning_estimate"},{"model_id":"qwen3.6-35b-a3b","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"BF16","estimated_output_tps":140,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"local-64gb","fit_status":"unsupported","reason":"Quantized weights exceed usable memory."},{"model_id":"kimi-k2.6-1t","hardware_id":"local-128gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":16,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"cloud-h200","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":36,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-flash-open","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":45,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-h200","fit_status":"unsupported","reason":"Single-device memory is insufficient at the tracked quantization."},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-b200","fit_status":"supported","quant":"Q2_K","estimated_output_tps":14,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"local-64gb","fit_status":"supported","quant":"BF16","estimated_output_tps":38,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"cloud-h200","fit_status":"supported","quant":"BF16","estimated_output_tps":80,"evidence":"planning_estimate"}],"decision_support":{"recommended_stack":[{"route":"hardest_long_horizon","primary":"fable-5","effort":"high","fallback":"gpt-5.6-sol","why":"Use Fable where long-horizon reliability warrants cost and retention policy is acceptable."},{"route":"default_engineering","primary":"gpt-5.6-terra","effort":"medium","fallback":"gpt-5.6-sol","why":"Best balanced default in the current benchmark/cost set."},{"route":"high_volume_subagents","primary":"gpt-5.6-luna","effort":"low","fallback":"gemini-3.5-flash-lite","why":"Luna retains stronger coding-agent evidence; Flash-Lite is the throughput-first fallback for extraction and delegated subtasks."},{"route":"privacy_local","primary":"qwen3.6-35b-a3b","effort":null,"fallback":"deepseek-v4-flash-open","why":"Local-first route where external retention is unacceptable."},{"route":"vision_long_context","primary":"gemini-3.1-pro","effort":"high","fallback":"gpt-5.6-sol","why":"Prefer for multimodal context; validate preview stability before production."}],"routing_rules":["Route by task risk and evidence, not provider family.","Start bulk work on Luna or Terra and escalate only after an explicit verification failure.","Do not send ZDR-required data to Fable 5 because the model requires 30-day retention.","Record model, effort, cache state, region and harness version for every internal speed comparison.","Use self-hosted routes only when the selected hardware row is supported; never infer performance from an empty cell."]},"actions":[{"priority":"P0","owner":"Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","status":"open","order":1},{"priority":"P0","owner":"Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","status":"open","order":2},{"priority":"P0","owner":"Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","status":"open","order":3},{"priority":"P1","owner":"Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","status":"open","order":4},{"priority":"P1","owner":"Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","status":"open","order":5},{"priority":"P2","owner":"Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics.","status":"open","order":6}],"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the price-change history using OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. Added the paid Codex and ChatGPT Work credit reduction, unchanged subscription prices and quota budgets, and Sol API Fast mode (up to 2.5× Standard speed at 2× price)."},{"date":"2026-08-01","tag":"pricing","text":"Refreshed OpenAI's live GPT-5.6 API rate card: Terra is now $2/$12 per million input/output tokens and Luna is $0.20/$1.20, with cached input at $0.20 and $0.02 respectively. Recomputed the 30/70 workload blend, cost-efficiency scores, SCQ compounds and burn baselines. Superseded on August 7 with the official July 30 effective date and full change details."},{"date":"2026-08-01","tag":"benchmark","text":"Recorded DeepSeek V4-Flash-0731's July 31 public-beta evidence: Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in DeepSeek Harness minimal mode. Additional vendor-reported NL2Repo, Cybergym, Toolathlon and Automation Bench results remain narrative evidence rather than new shared columns."},{"date":"2026-07-22","tag":"method","text":"Marked the open-weight quality history as a Fable-anchored source series that is rebased in the browser against the selected reference; fixed task-speed remains explicitly Fable-indexed."},{"date":"2026-07-22","tag":"model","text":"Added Gemini 3.6 Flash and Gemini 3.5 Flash-Lite benchmark, pricing, context and throughput evidence from Google and Artificial Analysis; retained missing task-time fields as null."},{"date":"2026-07-21","tag":"model","text":"Added official Kimi K3 API pricing and derived the documented 30/70 workload blend and cost-efficiency score; no task-speed compound is shown because comparable speed evidence remains unavailable."},{"date":"2026-07-19","text":"Added a dated GPU-hosting price tracker with billing-basis normalization, historical Runpod price changes, current Contabo configurations, and quote-aware Infomaniak options. Replaced stale Vast.ai constants with live-marketplace status."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, context, benchmark records and reference eligibility."},{"date":"2026-07-18","tag":"method","text":"Introduced documented quality, speed, cost, SCQ and six-axis capability composites with explicit missing-data rules."},{"date":"2026-07-18","tag":"hardware","text":"Replaced blank hardware fit speeds with supported/unsupported states and evidence-labelled planning estimates."},{"date":"2026-07-18","tag":"routing","text":"Updated recommended routing to Fable for hardest retained-data work, Terra for default engineering and Luna for high-volume subagents."}],"sources":[{"id":"openai-gpt-5-6","title":"GPT-5.6 launch and evaluations","url":"https://openai.com/index/gpt-5-6/","accessed":"2026-07-18"},{"id":"moonshot-kimi-k3-pricing","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing","accessed":"2026-07-21"},{"id":"openai-models","title":"OpenAI model catalog","url":"https://developers.openai.com/api/docs/models","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-terra","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-luna","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-price-performance","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","accessed":"2026-08-07"},{"id":"openai-gpt-5-6-chat-access","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","accessed":"2026-08-07"},{"id":"deepseek-v4-flash-0731","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","accessed":"2026-08-01"},{"id":"deepseek-v4-pricing","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/","accessed":"2026-08-01"},{"id":"anthropic-models","title":"Claude models overview","url":"https://platform.claude.com/docs/en/about-claude/models/overview","accessed":"2026-07-18"},{"id":"anthropic-retention","title":"Claude API and data retention","url":"https://platform.claude.com/docs/en/manage-claude/api-and-data-retention","accessed":"2026-07-18"},{"id":"google-gemini-3-5","title":"Gemini 3.5 Flash","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash","accessed":"2026-07-18"},{"id":"google-gemini-3-6","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash","accessed":"2026-07-22"},{"id":"google-gemini-3-5-flash-lite","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite","accessed":"2026-07-22"},{"id":"google-flash-july-2026","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/","accessed":"2026-07-22"},{"id":"aa-gemini-3-6-flash","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash","accessed":"2026-07-22"},{"id":"aa-gemini-3-5-flash-lite","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","accessed":"2026-07-22"},{"id":"moonshot-kimi-k3","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3","accessed":"2026-07-18"},{"id":"moonshot-models","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models","accessed":"2026-07-18"},{"id":"openai-gpt-5-5","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/","accessed":"2026-07-18"}],"hosting_prices":{"as_of":"2026-07-19","hours_per_month":730,"methodology":"Preserve provider currency and billing basis. normalized_hourly is a comparison-only derivation for fixed monthly plans; monthly_equivalent multiplies hourly rates by 730. It excludes tax, storage, egress, public IPs, support, discounts, utilization, and model throughput. Append observations instead of replacing them.","trend_policy":"Track the same provider, GPU, service tier, region basis, and currency. A changed SKU starts a new series. Two or more observations produce a trend; one observation is a dated baseline.","offers":[{"id":"runpod-h100-sxm-secure","provider":"Runpod","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$2.99/hr","normalized_hourly":2.99,"monthly_equivalent":2182.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":3.99,"source_n":64},{"date":"2026-07-19","value":2.99,"source_n":63}]},{"id":"runpod-a100-sxm-secure","provider":"Runpod","configuration":"1 x A100 SXM","gpu":"A100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$1.49/hr","normalized_hourly":1.49,"monthly_equivalent":1087.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.94,"source_n":64},{"date":"2026-07-19","value":1.49,"source_n":63}]},{"id":"runpod-l40s-secure","provider":"Runpod","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$0.99/hr","normalized_hourly":0.99,"monthly_equivalent":722.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.19,"source_n":64},{"date":"2026-07-19","value":0.99,"source_n":63}]},{"id":"hyperstack-h100-sxm","provider":"Hyperstack","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$2.40/hr","normalized_hourly":2.4,"monthly_equivalent":1752,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":2.4,"source_n":68}]},{"id":"hyperstack-h200-sxm","provider":"Hyperstack","configuration":"1 x H200 SXM","gpu":"H200 SXM","gpu_count":1,"vram_gb":141,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$3.50/hr","normalized_hourly":3.5,"monthly_equivalent":2555,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":3.5,"source_n":68}]},{"id":"contabo-l40s-monthly","provider":"Contabo","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€751/mo","normalized_hourly":1.0288,"monthly_equivalent":751,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":1.0288,"source_n":65}]},{"id":"contabo-h100-monthly","provider":"Contabo","configuration":"1 x H100","gpu":"H100","gpu_count":1,"vram_gb":80,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€1,838/mo","normalized_hourly":2.5178,"monthly_equivalent":1838,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.5178,"source_n":65}]},{"id":"contabo-h200-nvl-monthly","provider":"Contabo","configuration":"1 x H200 NVL","gpu":"H200 NVL","gpu_count":1,"vram_gb":141,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€2,149/mo","normalized_hourly":2.9438,"monthly_equivalent":2149,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.9438,"source_n":65}]},{"id":"infomaniak-l40s","provider":"Infomaniak","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Switzerland","billing":"Usage-based Public Cloud","price_basis":"Live calculator / availability validation","currency":"CHF","current_price_label":"Calculator / validate","normalized_hourly":null,"monthly_equivalent":null,"price_status":"not_publicly_exposed","checked":"2026-07-19","source_n":66,"history":[]},{"id":"vast-a6000-market","provider":"Vast.ai","configuration":"1 x RTX A6000","gpu":"RTX A6000","gpu_count":1,"vram_gb":48,"region":"Marketplace","billing":"Per second; on-demand, reserved, or interruptible","price_basis":"Live host marketplace","currency":"USD","current_price_label":"Live marketplace","normalized_hourly":null,"monthly_equivalent":null,"price_status":"dynamic_marketplace","checked":"2026-07-19","source_n":69,"history":[]}]}},"model_roster":{"schema_version":"2.0-preview","researched_at":"2026-08-07","default_reference":"fable-5","inclusion_policy":{"default_limit_per_provider":3,"default_scope":"core","summary":"Show a flagship, a balanced or fast model, and one distinctive specialist per provider. Keep siblings and emerging models in expandable extended and watchlist scopes.","promotion_rule":"Promote a watchlist family when at least two signals hold: current official release, meaningful adoption or discussion, differentiated capability or efficiency, and reproducible availability."},"speed_methodology":{"dimensions":["time_to_first_token","output_tokens_per_second","end_to_end_task_time"],"rule":"Compare speed only at the exact model, configuration, provider, region, date, workload, and cache condition. Vendor claims are labelled and unknown values remain unknown.","bands":{"fast":"Designed or measured for low interaction latency","balanced":"General-purpose latency and quality trade-off","deliberate":"Higher reasoning depth or multi-agent execution","unknown":"No comparable evidence"}},"models":[{"id":"fable-5","name":"Claude Fable 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","reference":true,"context":"1M","price":"$10 / $50","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"deliberate","speed_note":"Anthropic comparative latency: slower. No Fable fast mode.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"30-day retention; no zero-data-retention option","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"opus-5","name":"Claude Opus 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Anthropic comparative latency: moderate. Research-preview fast mode claims up to 2.5x higher output throughput at $10 / $50 per MTok.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"Adaptive thinking is on by default; disabling it is supported only through high effort. Independent benchmark results vary materially by effort level and should not be treated as a single generic score.","source":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5"},{"id":"opus-4.8","name":"Claude Opus 4.8","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Optional API fast mode claims up to 2.5x output throughput at premium pricing.","availability":"GA","source":"https://platform.claude.com/docs/en/build-with-claude/fast-mode"},{"id":"sonnet-5","name":"Claude Sonnet 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$3 / $15","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"fast","speed_note":"Anthropic comparative latency: fast.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"haiku-4.5","name":"Claude Haiku 4.5","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"200K","price":"$1 / $5","license":"Proprietary","control":"fixed","levels":[],"default_level":"n/a","speed_band":"fast","speed_note":"Anthropic’s latest verified Haiku and fastest listed Claude; no Haiku 5 is in the current official catalog.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports max reasoning completes comparable AAII work in 61% less time than Fable 5. API Fast mode, introduced July 30, claims up to 2.5x Standard speed at 2x price with unchanged intelligence; no comparable output-tokens/s figure is published.","availability":"GA; ChatGPT, Codex, API","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$2 / $12; cached input $0.20","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 20% from $2.50/$15; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; subscription prices and quota budgets unchanged","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$0.20 / $1.20; cached input $0.02","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"fast","speed_note":"OpenAI positions Luna as the fastest tier and reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 80% from $1/$6; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; rolling out as Free and Go default with unlimited text chats and a Think option","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.5","name":"GPT-5.5","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M API / 400K Codex","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports GPT-5.4-class per-token latency; Codex fast mode is 1.5x throughput at 2.5x cost.","availability":"API, ChatGPT and Codex","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"extended","status":"stable","context":"1M","price":"$30 / $180","license":"Proprietary","control":"reasoning effort","levels":["medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Higher-accuracy tier for difficult work; no comparable public output-throughput figure.","availability":"API and eligible ChatGPT plans","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"harness_bundled","scope":"extended","status":"stable","context":"400K","price":"Included by plan / routed service","license":"Proprietary","control":"fixed","levels":[],"default_level":"provider default","speed_band":"fast","speed_note":"Low-latency ChatGPT route; do not equate service behavior with the reasoning model API.","availability":"ChatGPT","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"preview","context":"1M","price":"$2 / $12","license":"Proprietary","control":"thinking level","levels":["low","medium","high"],"default_level":"high","speed_band":"deliberate","speed_note":"Google warns high thinking may significantly delay the first answer token.","availability":"Preview","source":"https://ai.google.dev/gemini-api/docs/thinking"},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$1.50 / $7.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Artificial Analysis measured about 304 output tok/s at high thinking; Google reports fewer reasoning turns and tool calls than 3.5 Flash.","availability":"GA via Gemini API, AI Studio, Gemini app and Antigravity","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$1.50 / $9","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Still offered; migrate new general Flash workloads to 3.6 Flash for lower output price and stronger agentic performance.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash"},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$0.30 / $2.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Artificial Analysis measured about 490 output tok/s; optimized for high-volume subagents, document parsing and extraction.","availability":"GA via Gemini API, AI Studio and Gemini app rollout","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$0.25 / $1.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Still offered; Gemini 3.5 Flash-Lite is the current high-throughput migration target.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite"},{"id":"grok-4.5","name":"Grok 4.5","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"500K","price":"$2 / $6","license":"Proprietary","control":"reasoning effort","levels":["low","medium","high"],"default_level":"high","speed_band":"fast","speed_note":"xAI reports 80 output tokens/s; vendor measurement conditions are not fully comparable here.","availability":"API and integrations; verify regional availability","source":"https://x.ai/news/grok-4-5"},{"id":"grok-4.20-multi-agent","name":"Grok 4.20 Multi-Agent","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"extended","status":"specialist","context":"1M","price":"$1.25 / $2.50","license":"Proprietary","control":"agent count","levels":["low","medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Levels control a 4- or 16-agent process, not ordinary reasoning depth.","availability":"API","source":"https://docs.x.ai/developers/model-capabilities/text/multi-agent"},{"id":"composer-2.5","name":"Composer 2.5","family":"Composer","provider":"Cursor","region":"US","channel":"harness_bundled","scope":"core","status":"stable","context":"Managed by Cursor","price":"$0.50 / $2.50 standard","license":"Proprietary","control":"speed variant","levels":["standard","fast"],"default_level":"fast","speed_band":"fast","speed_note":"Fast is the default; Cursor documents no public low/medium/high effort selector.","availability":"Cursor only","source":"https://cursor.com/blog/composer-2-5"},{"id":"kimi-k3","name":"Kimi K3","family":"Kimi K3","provider":"Moonshot AI","region":"China","channel":"commercial_api_open_weights_announced","scope":"core","status":"stable","context":"1M","price":"$3 / $15; cached input $0.30","license":"Terms pending weight release","control":"reasoning effort","levels":["max"],"default_level":"max","speed_band":"deliberate","speed_note":"No comparable K3 output-rate figure at launch; max thinking is always enabled.","availability":"API, Kimi, Kimi Work and Kimi Code; weights promised by July 27","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"id":"kimi-k2.7-code","name":"Kimi K2.7 Code","family":"Kimi K2.7","provider":"Moonshot AI","region":"China","channel":"commercial_api","scope":"core","status":"specialist","context":"256K","price":"Current API pricing","license":"Proprietary API","control":"service variant","levels":["standard","high-speed"],"default_level":"standard","speed_band":"fast","speed_note":"High-speed service is documented at about 180 tok/s and up to 260 tok/s for short contexts.","availability":"Kimi API and Kimi Code","source":"https://platform.kimi.ai/docs/models"},{"id":"deepseek-v4","name":"DeepSeek V4","family":"DeepSeek V4","provider":"DeepSeek","region":"China","channel":"open_weight","scope":"core","status":"public_beta","context":"1M","price":"Flash $0.14 / $0.28; cached input $0.0028","license":"MIT","control":"mode","levels":["non-think","think","max"],"default_level":"think","speed_band":"balanced","speed_note":"V4-Flash-0731 reports 82.7 on Terminal-Bench 2.1 at max effort in DeepSeek Harness minimal mode; comparable latency remains unknown.","availability":"V4-Flash-0731 API public beta and weights; future 2x peak-hours pricing announced without an effective date","variants":["Pro","Flash"],"source":"https://api-docs.deepseek.com/updates/"},{"id":"qwen-3.6","name":"Qwen3.6","family":"Qwen3.6","provider":"Alibaba Qwen","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"mode and token budget","levels":["non-thinking","thinking","budget"],"default_level":"thinking","speed_band":"balanced","speed_note":"Budget is provider-native; do not translate it into invented effort labels.","availability":"API and weights","variants":["27B","35B-A3B"],"source":"https://github.com/QwenLM/Qwen3.6"},{"id":"glm-5.2","name":"GLM-5.2","family":"GLM-5","provider":"Z.ai","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"MIT","control":"effort","levels":["adaptive"],"default_level":"adaptive","speed_band":"unknown","speed_note":"Architecture efficiency claims are not a comparable API latency measurement.","availability":"API and weights","source":"https://z.ai/blog/glm-5.2"},{"id":"kimi-k2.5","name":"Kimi K2.5","family":"Kimi K2","provider":"Moonshot AI","region":"China","channel":"open_weight","scope":"extended","status":"deprecated","context":"256K","price":"API and self-hosted","license":"Modified MIT","control":"mode","levels":["instant","thinking"],"default_level":"thinking","speed_band":"balanced","speed_note":"No comparable official throughput measurement; Instant is the lower-latency mode.","availability":"Existing users only; platform sunset August 31, 2026","source":"https://platform.kimi.ai/docs/models"},{"id":"minimax-m3","name":"MiniMax M3","family":"MiniMax M3","provider":"MiniMax","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"Community; non-commercial by default","control":"thinking mode","levels":["disabled","adaptive","enabled"],"default_level":"adaptive","speed_band":"balanced","speed_note":"Vendor claims decode improvements versus M2, not cross-provider latency.","availability":"API and weights","caveat":"License is not permissive open source","source":"https://www.minimax.io/blog/minimax-m3"},{"id":"mistral-small-4","name":"Mistral Small 4","family":"Mistral 3/4","provider":"Mistral AI","region":"EU","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"reasoning effort","levels":["none","high"],"default_level":"none","speed_band":"fast","speed_note":"Vendor claims 40% lower completion time and 3x requests/s versus Small 3.","availability":"API and weights","source":"https://mistral.ai/it/news/mistral-small-4/"},{"id":"llama-4","name":"Llama 4","family":"Llama 4","provider":"Meta","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"10M advertised for Scout","price":"Self-hosted","license":"Llama 4 Community License","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Performance depends on deployment; no comparable official API latency.","availability":"Weights","variants":["Scout","Maverick"],"caveat":"Not OSI-open; EU multimodal and large-platform restrictions apply","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"id":"gemma-3","name":"Gemma 3","family":"Gemma 3","provider":"Google","region":"US","channel":"open_weight","scope":"extended","status":"stable","context":"128K","price":"Self-hosted","license":"Google Gemma Terms","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Deployment dependent; optimized for edge and local use.","availability":"Weights","variants":["1B","4B","12B","27B"],"source":"https://ai.google.dev/gemma/docs/core/model_card_3"},{"id":"gpt-oss","name":"gpt-oss","family":"gpt-oss","provider":"OpenAI","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"128K","price":"Self-hosted","license":"Apache-2.0 plus usage policy","control":"reasoning effort","levels":["low","medium","high"],"default_level":"medium","speed_band":"unknown","speed_note":"Depends on hardware and serving stack.","availability":"Weights","variants":["120b","20b"],"source":"https://openai.com/index/introducing-gpt-oss/"},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","family":"Apertus v1.1","provider":"Swiss AI Initiative","region":"Switzerland","channel":"open_weight","scope":"core","status":"stable","context":"4K","price":"Self-hosted","license":"Apache-2.0","control":"sampling parameters","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"No comparable throughput measurement is published; performance depends on hardware, runtime and quantization.","availability":"Weights; Transformers, vLLM, SGLang and quantized checkpoints","variants":["0.5B","1.5B","4B"],"source":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"id":"inkling","name":"Inkling","family":"Inkling","provider":"Thinking Machines Lab","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"No metered API; open weights + Tinker fine-tuning platform","license":"Apache 2.0","control":"n/a","levels":[],"default_level":null,"speed_band":"unknown","speed_note":"No published throughput figure; not yet independently benchmarked for this roster.","availability":"Open weights (Hugging Face), Tinker fine-tuning platform, third-party inference providers","source":"https://thinkingmachines.ai/model-card/inkling/"}],"watchlist":[{"name":"Step 3.5 Flash","provider":"StepFun","region":"China","license":"Apache-2.0","signal":"Official 100-300 output tok/s claim","source":"https://github.com/stepfun-ai/Step-3.5-Flash"},{"name":"MiMo-V2-Flash","provider":"Xiaomi","region":"China","license":"Apache-2.0","signal":"Thinking toggle and generation-speed claim","source":"https://github.com/XiaomiMiMo/MiMo-V2-Flash"},{"name":"Seed-OSS-36B","provider":"ByteDance","region":"China","license":"Apache-2.0","signal":"512K context and controllable thinking budget","source":"https://github.com/ByteDance-Seed/seed-oss"},{"name":"Hy3 Preview","provider":"Tencent","region":"China","license":"Tencent Hy Community License","signal":"Preview multimodal MoE family","source":"https://github.com/Tencent-Hunyuan/Hy3-preview"}]}}&lt;/script&gt;
&lt;section class="tab-panel" id="tab-self-hosting"&gt;
 &lt;h2 class="sect-head"&gt;Hardware × model fit&lt;/h2&gt;
 &lt;p class="sect-sub"&gt;Quantization, VRAM used, and tok/s estimates per hardware config. Quality % is vs selected reference.&lt;/p&gt;</description></item><item><title>Data Methodology</title><link>https://projectious-work.github.io/ai-market-research/docs/data-methodology/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/docs/data-methodology/</guid><description>&lt;p&gt;The dashboard has three canonical data inputs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;data/market-state.json&lt;/code&gt; preserves the original report schema and backward
compatibility.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;data/model-roster-v2.json&lt;/code&gt; contains the curated provider roster, model
configuration levels, availability, pricing labels, and speed evidence.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;data/report-metrics.json&lt;/code&gt; contains normalized metrics, chart encodings,
current economics, hardware fit, and decision-support data.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The build validates all three files and merges the latter two under
&lt;code&gt;marketData.model_roster&lt;/code&gt; and &lt;code&gt;marketData.report_metrics&lt;/code&gt;. Do not fetch data
at runtime: GitHub Pages remains a single, manually rebuilt static artifact.&lt;/p&gt;</description></item><item><title>04 Decisions</title><link>https://projectious-work.github.io/ai-market-research/report/04-decisions/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/report/04-decisions/</guid><description>&lt;div class="sr-report-scope"&gt;
&lt;script id="market-data" type="application/json"&gt;{"meta":{"generated_at":"2026-08-07T00:00:00Z","reference_default":"fable-5","report_metrics_file":"data/report-metrics.json"},"executive_summary":{"models":["**Claude Opus 5 is now generally available.** Anthropic's new `claude-opus-5` targets complex agentic coding and enterprise work with a 1M-token context window, 128K maximum output, adaptive thinking by default, and unchanged $5/$25 per MTok base pricing. Independent benchmark values are not yet recorded here.","**Google reset the Flash price-performance curve on July 21.** Gemini 3.6 Flash is GA at $1.50/$7.50 per million tokens with stronger coding and agentic results than 3.5 Flash, while Gemini 3.5 Flash-Lite reaches roughly 490 output tok/s at $0.30/$2.50 for high-volume subagents and extraction.","**Kimi K3 resets the open-weight frontier.** Moonshot reports a 2.8T sparse MoE, native vision, 1M context, and 99% of Fable 5 across 14 overlapping launch-suite evaluations; weights are promised by July 27.","**GPT-5.6 Luna and Terra became materially cheaper on July 30.** Luna fell 80% to $0.20/$1.20 per MTok and Terra 20% to $2/$12; their paid Codex and ChatGPT Work usage also consumes fewer credits, while subscription prices and quota budgets did not change.","**GPT-5.6 now spans Sol, Terra and Luna**, giving leaders a deliberate capability, balanced, and high-throughput ladder under one family. Sol API Fast mode replaces Priority Processing: OpenAI claims up to 2.5× Standard speed at 2× price, with unchanged intelligence.","**Luna is expanding beyond API routing.** OpenAI says it becomes the default for Free and Go users, with unlimited text chats and a higher-reasoning Think option rolling out subject to abuse guardrails. This is a ChatGPT product update, not an API capability change.","**Keep benchmark tables separate by harness.** OpenAI's GPT-5.6 launch results and Scale's public SWE-bench Pro leaderboard use different model versions and evaluation setups; use each as evidence, not as one directly rankable series.","**GPT-5.5 remains active in Codex and the API.** It is retained as a compatibility and portfolio option rather than being hidden by the newer 5.6 family.","**Claude Haiku 4.5 remains Anthropic’s latest verified Haiku.** No official Haiku 5 listing was found in Anthropic’s current model catalog."],"harnesses":["**Separate model capability from harness capability.** Tool execution, context management, isolation, and observability can dominate real workflow outcomes.","**Subscription access is not a production routing contract.** Validate API, credit-pool, and third-party harness policies before standardizing an operating model.","**Maintain at least one portable fallback path** across providers for high-value workflows and operational incidents.","**Gemini Managed Agents now add background execution, remote MCP, custom functions and credential refresh.** That makes Google a more credible managed-agent control plane, but it does not substitute for workload-specific evaluation."],"self_hosting":["**Open weights are now a strategic option, not only a cost play.** Kimi K3, DeepSeek, Qwen, Nemotron, Mistral, and Llama cover different sovereignty and specialization needs.","**Do not compare self-hosting at zero token cost.** Include accelerator rental or depreciation, power, utilization, serving staff, and measured throughput.","**Pilot against a defined workload and hardware envelope** before treating advertised context or parameter scale as deployable capacity."],"strategy":["**Run a portfolio, not a winner-takes-all model standard.** Reserve frontier reasoning for high-value decisions and route routine work to measured fast or efficient tiers.","**Instrument quality, latency, retries, and total workflow cost together.** Token price alone is not an operating metric.","**Review the portfolio quarterly and after major releases**, with explicit retirement, security, and fallback criteria."]},"headline_stats":[{"id":"frontier_count","label":"Models tracked","value":35,"unit":"","delta":"Current, fast and retained compatibility models","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_weight","label":"Open-weight","color":"#16866f"}],"history":[{"date":"2026-05-18","proprietary":20,"open_weight":3},{"date":"2026-07-22","proprietary":26,"open_weight":8},{"date":"2026-07-25","proprietary":27,"open_weight":8}],"note":"Release-tag roster snapshots; announced open-weight releases are counted with open-weight models."}},{"id":"best_open_pct","label":"Best open-weight vs selected reference","value":99,"unit":"%","delta":"Kimi K3; geometric mean across 14 overlapping vendor evals","delta_dir":"up","signal_id":"open_weight_quality"},{"id":"cheapest_frontier_api","label":"Lowest frontier blended API price","value":0.9,"unit":"$/Mtok","delta":"GPT-5.6 Luna; 30% input / 70% output blend","delta_dir":"down","signal_id":"blended_frontier_price"},{"id":"fastest_task_rate","label":"Fastest task-rate index","value":"3.2×","unit":"","delta":"Fable 5 = 1.0×; planning index, not tokens/second","delta_dir":"up","signal_id":"task_speed"},{"id":"max_context","label":"Largest usable context","value":"10M","unit":"tok","delta":"Llama 4 Scout; deployment constraints still apply","delta_dir":"up","trend":[{"date":"2026-01-01","value":1},{"date":"2026-04-01","value":1},{"date":"2026-07-01","value":10}]},{"id":"harness_count","label":"Harnesses tracked","value":16,"unit":"","delta":"Commercial and open agent environments","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_source","label":"Open source","color":"#3d78c5"}],"history":[{"date":"2026-05-18","proprietary":3,"open_source":11},{"date":"2026-07-22","proprietary":6,"open_source":10}],"note":"Release-tag roster snapshots classified from each harness license."}},{"id":"documented_output_speed","label":"Fastest documented API output","value":490,"unit":"tok/s","delta":"Gemini 3.5 Flash-Lite; Artificial Analysis first-party API measurement","delta_dir":"up","signal_id":"output_throughput"},{"id":"quality_coverage","label":"Models with quality evidence","value":"31/35","unit":"","delta":"Unknown remains unknown; no zero-value substitution","delta_dir":"neutral","trend":[{"date":"2026-01-01","value":18},{"date":"2026-04-01","value":24},{"date":"2026-07-01","value":31}]}],"trends":{"best_open_vs_opus":{"label":"Best open-weight % vs Opus 4.7","history":[{"date":"2025-11-01","value":68},{"date":"2025-12-01","value":72},{"date":"2026-01-01","value":78},{"date":"2026-02-01","value":80},{"date":"2026-03-01","value":83},{"date":"2026-04-01","value":86},{"date":"2026-05-01","value":88},{"date":"2026-06-01","value":90}]},"median_frontier_output_price":{"label":"Median frontier $/Mtok output","history":[{"date":"2025-11-01","value":18},{"date":"2025-12-01","value":17},{"date":"2026-01-01","value":15.5},{"date":"2026-02-01","value":14.5},{"date":"2026-03-01","value":13},{"date":"2026-04-01","value":12.5},{"date":"2026-05-01","value":12},{"date":"2026-06-01","value":11}]},"models_released_per_month":{"label":"Notable model releases per month","history":[{"date":"2025-11-01","value":3},{"date":"2025-12-01","value":4},{"date":"2026-01-01","value":5},{"date":"2026-02-01","value":6},{"date":"2026-03-01","value":4},{"date":"2026-04-01","value":7},{"date":"2026-05-01","value":5},{"date":"2026-06-01","value":8}]}},"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the GPT-5.6 pricing history from OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. The report now records lower paid Codex and ChatGPT Work credit consumption, unchanged subscription prices and quota budgets, Sol API Fast mode (up to 2.5× Standard speed at 2× price), and the August 6 ChatGPT Luna access expansion."},{"date":"2026-08-01","tag":"pricing","text":"Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to $0.20/$1.20 from OpenAI's live model pages. Recomputed the tracked 30/70 workload blends and effort-burn ratios. Superseded on August 7 with OpenAI's published July 30 effective date and full price-change details."},{"date":"2026-08-01","tag":"model","text":"Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output, thinking/non-thinking modes, $0.14/$0.28 per-million pricing, $0.0028 cached input, and vendor-reported agent results including Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2."},{"date":"2026-07-28","tag":"model","text":"Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B active) open-weight (Apache 2.0) multimodal MoE with 1M-token context, released 2026-07-15. No first-party per-token API pricing is published (monetized via the Tinker fine-tuning platform); benchmark scores are published on the model card but not yet normalized into this roster's comparable set, so quality/speed/cost fields are left unknown rather than estimated."},{"date":"2026-07-25","tag":"model","text":"Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive thinking by default, $5/$25 per MTok base pricing, and official cloud-platform availability. Added independent Artificial Analysis evidence (61 Intelligence Index at max effort; 52.3 output tok/s) with effort-specific caveats. Gemini 3.5 Flash Cyber remains limited to CodeMender government and trusted-partner pilots; GPT-Live and Muse Spark 1.1 remain non-API products, so none were added to the API roster. Claude Opus 4.7 Fast Mode was removed July 24."},{"date":"2026-07-24","tag":"benchmark","text":"Refreshed current-source evidence: added OpenAI's GPT-5.6 launch table, Google's Managed Agents update, and Scale's public SWE-bench Pro leaderboard. Clarified that vendor launch tables and the public leaderboard are not directly comparable because their model versions and harnesses differ."},{"date":"2026-07-22","tag":"fix","text":"Made the selected reference propagate through open-weight quality headlines, market-signal history, comparison headings, model analytics, self-hosting quality and capability market position. Replaced the model and harness inventory mini-lines with stacked proprietary/open category areas based on release-tag roster snapshots."},{"date":"2026-07-22","tag":"model","text":"Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M context, 64K output, current API pricing, Artificial Analysis intelligence and throughput measurements, and Google’s published coding and agentic benchmarks."},{"date":"2026-07-22","tag":"data","text":"Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to governments and trusted CodeMender partners; Gemini Omni Flash and Nano Banana 2 Lite remain specialized media models rather than general-purpose roster entries."},{"date":"2026-07-21","tag":"model","text":"Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint: fully open Apache 2.0 weights and data, 4K context, 1.7T-token distillation, 1,811 languages, and official BF16, FP8, NVFP4A16, INT3, INT4 and INT6 variants."},{"date":"2026-07-21","tag":"harness","text":"Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI 0.144.6, Gemini CLI 0.51.0, OpenCode 1.18.4, Cline 4.0.10, and Goose 1.43.0; updated repository star snapshots and notable release capabilities."},{"date":"2026-07-21","tag":"model","text":"Added Moonshot's official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached input, and $15/M output; the 30/70 workload blend is $11.40/M before reasoning-effort effects."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark registers and selectable reference configurations. Added documented quality, speed, cost and capability composites in data/report-metrics.json."},{"date":"2026-07-18","tag":"routing","text":"Updated the action queue and recommended routing: Fable for hardest retained-data workloads, Terra for default engineering, Luna for high-volume subagents, with explicit escalation rules."},{"date":"2026-06-06","tag":"fix","text":"Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K context, 1000+ tok/s on Cerebras) and Hermes Agent v0.16.0 (Nous Research, MIT, self-hosted multi-platform agent) — both were incorrectly removed in v1.5.0 sweep."},{"date":"2026-06-06","tag":"policy","text":"Dashboard market sweep v1.5.0: real-world re-grounding. Replaced fictional Mythos/GPT-5.5-Cyber rows with verified models; added Nvidia Nemotron coalition, Kimi K2.6, GLM-5, Cohere Command A+, SubQ 1M-Preview."},{"date":"2026-06-04","tag":"model","text":"Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context, NVIDIA Open Model License) at Computex — first frontier-scale open model from Nvidia."},{"date":"2026-06-04","tag":"model","text":"Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral, Perplexity, Reflection AI, Sarvam, Thinking Machines Lab as inaugural members."},{"date":"2026-06-01","tag":"model","text":"Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model."}],"actions":["P0 · Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","P0 · Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","P0 · Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","P1 · Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","P1 · Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","P2 · Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics."],"models":[{"id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","tier":"frontier","released":"2026-06-09","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":80,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii":59.9,"coding_agent_index":77.2,"deep_swe":69.7,"terminal_bench":83.1,"agents_last_exam":40.5,"api_in":10,"api_out":50,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":52.3,"subscription":"Claude API and supported cloud platforms","notes":"Default reference for v2. Anthropic describes Fable 5 as its most capable widely released model. Adaptive thinking is always on. Comparative latency is slower. Retention is 30 days and zero-data-retention is not available.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"status":"stable","speed_class":"deliberate","speed_evidence":"vendor-qualitative","capability_levels":{"coding":49,"reasoning":50,"knowledge":48,"comms":48,"multimodal":45,"agentic":50}},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":64.6,"swe_verified":null,"aaii":58.9,"coding_agent_index":80,"deep_swe":72.7,"terminal_bench":88.8,"agents_last_exam":52.7,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Flagship GPT-5.6 tier. API Fast mode replaced Priority Processing on July 30: OpenAI claims up to 2.5× Standard speed at 2× Standard price with no intelligence change. Speed is otherwise stored as an end-to-end task index in report-metrics.json, not tok/s.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":49,"knowledge":48,"comms":48,"multimodal":48,"agentic":49}},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":63.4,"swe_verified":null,"aaii":55,"coding_agent_index":77.4,"deep_swe":69.6,"terminal_bench":87.4,"agents_last_exam":50.4,"api_in":2,"api_out":12,"api_cache_hit":0.2,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Balanced GPT-5.6 tier and recommended default engineering route. On July 30, API list pricing fell 20% from $2.50/$15 to $2/$12 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":48,"reasoning":46,"knowledge":46,"comms":46,"multimodal":45,"agentic":48}},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","tier":"fast","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":62.7,"swe_verified":null,"aaii":51.2,"coding_agent_index":74.6,"deep_swe":67.2,"terminal_bench":84.7,"agents_last_exam":50.3,"api_in":0.2,"api_out":1.2,"api_cache_hit":0.02,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Fastest, most affordable GPT-5.6 tier; recommended for high-volume subagents with verification and escalation. On July 30, API list pricing fell 80% from $1/$6 to $0.20/$1.20 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change. ChatGPT is also rolling Luna out as the Free and Go default, with unlimited text chats and a Think option; that product change does not alter API routing.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":46,"reasoning":44,"knowledge":43,"comms":45,"multimodal":43,"agentic":46}},{"id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","tier":"frontier","released":"2026-07-16","license":"open weights announced; terms pending weight release","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":null,"tok_per_sec":null,"subscription":"Kimi API, Kimi Code, Kimi Work","reasoning_capable":true,"effort_default":"max","effort_levels":["max"],"notes":"2.8T sparse MoE with 16/896 experts active, native vision and always-on thinking. API pricing is $3/M uncached input, $0.30/M cached input and $15/M output. Full weights promised by July 27, 2026.","quality_vs_fable":99.2,"quality_evidence":"Geometric mean across 14 overlapping values in Moonshot’s launch comparison; vendor-reported, max/xhigh settings.","capability_levels":{"coding":47,"reasoning":46,"knowledge":47,"comms":46,"multimodal":47,"agentic":48}},{"id":"opus-5","name":"Claude Opus 5","provider":"Anthropic","tier":"frontier","released":"2026-07-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"Claude API, Bedrock, Google Cloud, Microsoft Foundry","notes":"Current Anthropic Opus generation. 1M context and 128K maximum output. Adaptive thinking is enabled by default; effort defaults to high. Artificial Analysis reports a 61 Intelligence Index and 52.3 output tok/s at max effort; these figures are effort-specific. Research-preview Fast Mode is Claude API-only at $10/$50 per MTok and claims up to 2.5x higher output throughput.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"speed_class":"balanced","speed_evidence":"vendor-qualitative","cite":[86,87]},{"id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","tier":"frontier","released":"2026-05-28","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":69.2,"swe_verified":88.6,"livecodebench":82,"aime":90,"tau2":86,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Max 20× $200/mo · Max 5× $100/mo","notes":"Previous Opus generation, retained as an active compatibility option. SWE-V 88.6%, SWE-Pro 69.2%, AAII 61.4. Fast Mode reduced 3× to $10/$50 (was $30/$150 on 4.7). 1M context standard.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":50,"comms":50,"multimodal":34,"agentic":50}},{"id":"opus-4.7","name":"Claude Opus 4.7","provider":"Anthropic","tier":"frontier","released":"2026-04-16","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":64.3,"swe_verified":87.6,"livecodebench":79,"aime":88,"tau2":84,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":50,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Now legacy as of Opus 4.8 release May 28. Pricing unchanged. SWE-Verified 87.6%, SWE-Pro 64.3%. Fast Mode was removed July 24, 2026; standard-speed API access remains active.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":49,"comms":50,"multimodal":32,"agentic":50}},{"id":"sonnet-4.6","name":"Claude Sonnet 4.6","provider":"Anthropic","tier":"frontier","released":"2026-02-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":44,"swe_verified":77,"livecodebench":76,"aime":85,"tau2":81,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":50,"tok_per_sec":80,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Best code style/intent understanding. With cache+batch: $0.30/$7.50 effective.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","max"],"cache_discount":0.1,"capability_levels":{"coding":41,"reasoning":40,"knowledge":39,"comms":42,"multimodal":31,"agentic":41}},{"id":"haiku-4.5","name":"Claude Haiku 4.5","provider":"Anthropic","tier":"fast","released":"2025-10-15","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":28,"swe_verified":62,"livecodebench":58,"aime":70,"tau2":65,"api_in":1,"api_out":5,"api_cache_hit":0.1,"batch_discount":50,"tok_per_sec":110,"subscription":"Available in all tiers","notes":"5× cheaper than Sonnet. Triage/classification champion.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.1,"capability_levels":{"coding":28,"reasoning":18,"knowledge":29,"comms":30,"multimodal":18,"agentic":28}},{"id":"gpt-5.5","name":"GPT-5.5","provider":"OpenAI","tier":"frontier","released":"2026-04-23","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M (272K threshold)","swe_pro":58.6,"swe_verified":88.7,"livecodebench":84,"aime":92,"tau2":82,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Pro $200 · Pro Lite $100 · Plus $20","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"OpenAI flagship; default in ChatGPT (Instant variant since May 5, 2026). 1M context with 2× input/1.5× output surcharge above 272K. Reasoning tokens billed as output.","cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":41,"knowledge":49,"comms":40,"multimodal":41,"agentic":38}},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","provider":"OpenAI","tier":"frontier","released":"2026-04-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":60,"swe_verified":90,"livecodebench":86,"aime":94,"tau2":84,"api_in":30,"api_out":180,"api_cache_hit":3,"batch_discount":50,"tok_per_sec":35,"subscription":"Pro $200 only","reasoning_capable":true,"effort_default":"high","effort_levels":["medium","high","xhigh"],"notes":"Highest-stakes reasoning tier. $30/$180. Available in ChatGPT Pro $200 and as API. 6× cost of base 5.5.","cache_discount":0.25,"capability_levels":{"coding":47,"reasoning":48,"knowledge":50,"comms":42,"multimodal":42,"agentic":42}},{"id":"gpt-5.4","name":"GPT-5.4","provider":"OpenAI","tier":"frontier","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":57.7,"swe_verified":81,"livecodebench":81,"aime":90,"tau2":80,"api_in":2.5,"api_out":15,"api_cache_hit":0.25,"batch_discount":50,"tok_per_sec":60,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Best quality-per-credit. Held SWE-Pro lead Feb-April. Default for everyday coding.","cache_discount":0.1,"capability_levels":{"coding":40,"reasoning":40,"knowledge":40,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-mini","name":"GPT-5.4 Mini","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":38,"swe_verified":73,"livecodebench":71,"aime":78,"tau2":70,"api_in":0.4,"api_out":1.6,"api_cache_hit":0.04,"batch_discount":50,"tok_per_sec":130,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"94% of GPT-5.4 coding at 6× less. Best subagent. ~1/20 burn vs GPT-5.5. Collapses at 64K+ context.","cache_discount":0.1,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-nano","name":"GPT-5.4 Nano","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":25,"swe_verified":60,"livecodebench":58,"aime":65,"tau2":55,"api_in":0.1,"api_out":0.4,"api_cache_hit":0.01,"batch_discount":50,"tok_per_sec":200,"subscription":"API only","reasoning_capable":true,"effort_default":"low","effort_levels":["minimal","low","medium"],"notes":"Smallest reasoning model. API-only. For embeddable/edge inference at near-zero cost.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":20,"comms":20,"multimodal":20,"agentic":20}},{"id":"gpt-5.3-codex","name":"GPT-5.3-Codex","provider":"OpenAI","tier":"frontier","released":"2026-01-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":56.8,"swe_verified":85,"livecodebench":82,"aime":87,"tau2":78,"api_in":1.5,"api_out":10,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":70,"subscription":"All paid tiers · Code Review uses this","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Coding-specialised. ~⅓ burn vs GPT-5.5 for ~2pts less SWE-Pro. Often best $/quality for execution turns.","cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":28,"knowledge":28,"comms":27,"multimodal":8,"agentic":38}},{"id":"gpt-5.3-codex-spark","name":"GPT-5.3-Codex-Spark","provider":"OpenAI","tier":"fast","released":"2026-02-12","license":"proprietary","jurisdiction":"US","context":128000,"context_label":"128K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":1000,"subscription":"ChatGPT Pro — research preview only","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"Smaller, latency-first sibling of GPT-5.3-Codex. Released Feb 12, 2026. 1000+ tok/s on Cerebras hardware. ChatGPT Pro research preview only — not in API at launch; separate preview rate-limit pool (no standard credit burn). Text-only. Target use: real-time micro-edits, live pair-programming in Codex app/CLI/VS Code.","cache_discount":null,"capability_levels":{"coding":38,"reasoning":22,"knowledge":22,"comms":25,"multimodal":0,"agentic":28}},{"id":"gpt-5.2","name":"GPT-5.2","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":52,"swe_verified":78,"livecodebench":76,"aime":84,"tau2":75,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Available but not recommended","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Codex picker keeps it for 'long-running agents' (specifically tuned for autonomy). Otherwise eclipsed by 5.3-Codex.","cache_discount":0.1},{"id":"gpt-5.2-codex","name":"GPT-5.2-Codex","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":50,"swe_verified":76,"livecodebench":75,"aime":82,"tau2":73,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Legacy","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Predecessor to 5.3-Codex. Legacy.","cache_discount":0.1},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","provider":"OpenAI","tier":"fast","released":"2026-05-05","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":35,"swe_verified":76,"livecodebench":72,"aime":81,"tau2":70,"api_in":1.5,"api_out":6,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":200,"subscription":"Default ChatGPT model · API chat-latest","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"ChatGPT default since May 5, 2026. 52.5% fewer hallucinations vs 5.3, 30% shorter responses, first Instant-class High-capability rating on cybersec/bio-chem.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":40,"comms":30,"multimodal":30,"agentic":10}},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","provider":"Google","tier":"frontier","released":"2026-02-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":54.2,"swe_verified":78,"livecodebench":76,"aime":85,"tau2":78,"api_in":2,"api_out":12,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":119,"subscription":"Gemini Advanced $20 · Ultra $100","notes":"Released Feb 19, 2026 (preview). Pricing doubles above 200K input tokens. 50% batch discount.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":39,"reasoning":41,"knowledge":49,"comms":41,"multimodal":50,"agentic":29}},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":58.7,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78,"agents_last_exam":null,"quality_vs_fable":81.6,"api_in":1.5,"api_out":7.5,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":304,"subscription":"Gemini API · AI Studio · Gemini app · Antigravity","notes":"GA stable ID gemini-3.6-flash. Google reports fewer tool calls and 17% fewer output tokens than 3.5 Flash on the AA Index workload; Computer Use is preview. Artificial Analysis measured about 304 output tok/s at high thinking.","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":42,"reasoning":38,"knowledge":45,"comms":40,"multimodal":47,"agentic":42}},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","provider":"Google","tier":"legacy","released":"2026-05-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55.1,"swe_verified":78,"livecodebench":70,"aime":75,"tau2":68,"aaii_v4_1":50,"swe_bench_pro":55.1,"deep_swe_v1_1":37,"terminal_bench_2_1":76.2,"agents_last_exam":null,"quality_vs_fable":76.1,"api_in":1.5,"api_out":9,"api_cache_hit":0.375,"batch_discount":50,"tok_per_sec":165,"subscription":"Free CLI: 1000 req/day at 1M context","notes":"Shipped GA at Google I/O May 19, 2026. Still offered, but Gemini 3.6 Flash is the recommended migration target with stronger agentic results and lower output-token pricing.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":30,"reasoning":30,"knowledge":40,"comms":30,"multimodal":40,"agentic":20}},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":54.2,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":36,"swe_bench_pro":54.2,"deep_swe_v1_1":null,"terminal_bench_2_1":54,"agents_last_exam":null,"quality_vs_fable":65.9,"api_in":0.3,"api_out":2.5,"api_cache_hit":0.03,"batch_discount":50,"tok_per_sec":490,"subscription":"Gemini API · AI Studio · Gemini app rollout","notes":"GA stable ID gemini-3.5-flash-lite. Google positions it for high-volume subagents, document parsing and structured extraction; Artificial Analysis measured about 490 output tok/s. Computer Use availability differs across Google documentation and should be validated per API surface.","reasoning_capable":true,"effort_default":"minimal","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":32,"reasoning":28,"knowledge":34,"comms":32,"multimodal":38,"agentic":35}},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash Lite","provider":"Google","tier":"legacy","released":"2026-03-10","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":22,"swe_verified":65,"livecodebench":60,"aime":68,"tau2":58,"api_in":0.25,"api_out":1.5,"api_cache_hit":0.0625,"batch_discount":50,"tok_per_sec":220,"subscription":"Free CLI + AI Studio","notes":"Lite tier; 1M context retained.","reasoning_capable":true,"effort_default":"off","effort_levels":["off","low","medium"],"cache_discount":0.25,"capability_levels":{"coding":22,"reasoning":22,"knowledge":30,"comms":25,"multimodal":35,"agentic":15}},{"id":"deepseek-v4-pro","name":"DeepSeek V4-Pro","provider":"DeepSeek","tier":"frontier","released":"2026-04-24","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":55.4,"swe_verified":80.6,"livecodebench":93.5,"aime":88,"tau2":78,"api_in":0.435,"api_out":0.87,"api_cache_hit":0.043,"batch_discount":null,"tok_per_sec":60,"subscription":"API only","notes":"MIT-licensed. Permanent pricing May 22, 2026 — Opus-class quality at ~1/10 cost. 1.6T/49B MoE, 1M context.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.0083,"capability_levels":{"coding":47,"reasoning":40,"knowledge":38,"comms":30,"multimodal":8,"agentic":28}},{"id":"deepseek-v4-flash","name":"DeepSeek V4-Flash","provider":"DeepSeek","tier":"fast","released":"2026-07-31","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"tau2":72,"terminal_bench_2_1":82.7,"agents_last_exam":25.2,"api_in":0.14,"api_out":0.28,"api_cache_hit":0.0028,"batch_discount":null,"tok_per_sec":90,"subscription":"API + open weights","notes":"V4-Flash-0731 public beta; 284B/13B active MoE, 1M context and 384K max output. DeepSeek reports Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in its unreleased minimal harness; internal DSBench results are not normalized here. A future 2x peak-hours rate is announced, but no effective date is published.","reasoning_capable":true,"effort_default":"think","effort_levels":["non-think","think"],"cache_discount":0.02,"capability_levels":{"coding":40,"reasoning":30,"knowledge":30,"comms":25,"multimodal":5,"agentic":25}},{"id":"minimax-m2.7","name":"MiniMax M2.7","provider":"MiniMax","tier":"frontier","released":"2026-03-18","license":"open-weight","jurisdiction":"China","context":205000,"context_label":"205K","swe_pro":40,"swe_verified":74,"livecodebench":72,"aime":80,"tau2":73,"api_in":0.3,"api_out":1.2,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":70,"subscription":"API + open-weight","notes":"Current flagship reasoner; MoE 230B/10B active.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.52,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":20,"agentic":20}},{"id":"grok-4.3","name":"Grok 4.3","provider":"xAI","tier":"frontier","released":"2026-05-06","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":45,"swe_verified":80,"livecodebench":80,"aime":88,"tau2":76,"api_in":1.25,"api_out":2.5,"api_cache_hit":0.31,"batch_discount":null,"tok_per_sec":75,"subscription":"SuperGrok $30 · Heavy $300","notes":"Current xAI flagship — aggressive pricing for frontier tier. Hybrid reasoning. Multimodal text+image.","reasoning_capable":true,"effort_default":"reasoning-on","effort_levels":["off","on"],"cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":44,"knowledge":42,"comms":34,"multimodal":34,"agentic":34}},{"id":"grok-4.1-fast","name":"Grok 4.1 Fast","provider":"xAI","tier":"fast","released":"2026-03-20","license":"proprietary","jurisdiction":"US","context":2000000,"context_label":"2M","swe_pro":30,"swe_verified":70,"livecodebench":65,"aime":75,"tau2":65,"api_in":0.2,"api_out":0.5,"api_cache_hit":0.05,"batch_discount":null,"tok_per_sec":140,"subscription":"SuperGrok $30","notes":"Cheapest large-context model on market. 2M context, $0.05/Mtok cached input.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.25,"capability_levels":{"coding":28,"reasoning":28,"knowledge":30,"comms":26,"multimodal":26,"agentic":22}},{"id":"mistral-medium-3.5","name":"Mistral Medium 3.5","provider":"Mistral","tier":"frontier","released":"2026-04-29","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":42,"swe_verified":77.6,"livecodebench":75,"aime":75,"tau2":65,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":75,"subscription":"Le Chat Pro €20","notes":"EU jurisdiction. 128B dense, Apache 2.0. Strongest non-Chinese open-weight coding agent. Vibe agents (GitHub/Linear/Jira/Sentry integrations). Pricing not verified; Medium 3 was $0.40/$2.00.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":40,"reasoning":30,"knowledge":40,"comms":30,"multimodal":10,"agentic":20}},{"id":"mistral-large-3","name":"Mistral Large 3","provider":"Mistral","tier":"frontier","released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"tau2":70,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":65,"subscription":"API + Le Chat","notes":"Cheapest premium output price in market among Western frontier. 675B/41B MoE, Apache 2.0.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":38,"reasoning":36,"knowledge":40,"comms":36,"multimodal":20,"agentic":28}},{"id":"subq-1m-preview","name":"SubQ 1M-Preview","provider":"SubQ","tier":"frontier","released":"2026-05-15","license":"proprietary","jurisdiction":"US","context":12000000,"context_label":"12M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":1,"api_out":5,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":80,"subscription":"API preview","notes":"First commercial subquadratic (non-transformer) LLM. ~1/5 frontier cost on long-context tasks. Capability rating estimated — public benchmarks pending.","reasoning_capable":null,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":30,"reasoning":32,"knowledge":36,"comms":30,"multimodal":5,"agentic":28}},{"id":"nemotron-3-ultra","name":"Nvidia Nemotron 3 Ultra","provider":"Nvidia","tier":"frontier","released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"tau2":76,"api_in":1.5,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":300,"subscription":"build.nvidia.com (closed API tier) + open weights","notes":"550B/55B MoE, hybrid Mamba-Transformer. ~300 tok/s. Open weights also available — see self_hosting. Capability levels are best estimates pending independent benchmarks.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":42,"reasoning":42,"knowledge":42,"comms":34,"multimodal":10,"agentic":34}},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","provider":"Swiss AI Initiative","tier":"fast","released":"2026-06-15","license":"Apache 2.0","jurisdiction":"Switzerland","context":4096,"context_label":"4K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Self-hosted weights","notes":"Largest newly released Apertus v1.1 distilled checkpoint. Dense 4.6B storage / 3.8B compute parameters, trained on 1.7T tokens, supports 1,811 languages, and ships in BF16 plus server and Apple-oriented quantizations. No comparable coding-agent, throughput, or API-price evidence is published.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null},{"id":"inkling","name":"Inkling","provider":"Thinking Machines Lab","tier":"frontier","released":"2026-07-15","license":"Apache 2.0","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Open weights (Hugging Face); fine-tuning and inference via the Tinker platform and third-party providers","notes":"975B-parameter multimodal MoE (41B active), 66-layer decoder-only transformer, 1M-token context. Text/image/audio input, text-only output. No first-party per-token API pricing published -- Thinking Machines monetizes via the Tinker fine-tuning platform rather than metered inference. A smaller Inkling-Small (12B active) companion model was released alongside it. Benchmark results are published on the model card across reasoning, agentic, coding, factuality, vision, audio, and safety categories but are not yet normalized into this roster's comparable benchmark set.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null}],"subscriptions":[{"provider":"Anthropic","tier":"Pro","price_usd":20,"limits":"~45 messages / 5h on Sonnet · limited Opus","models":"Sonnet 4.6, Haiku 4.5, limited Opus 4.8/4.7","features":"Chat only (Claude Code removed April 2026). From June 15, 2026 split into Chat pool + Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 5×","price_usd":100,"limits":"5× Pro quotas · ~225 msg/5h Sonnet · expanded Opus","models":"Full Opus 4.8/4.7 · Sonnet 4.6 · Haiku 4.5","features":"Cache reads included flat-rate. From June 15, 2026: Chat pool + separate Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 20×","price_usd":200,"limits":"20× Pro quotas","models":"All","features":"For heavy Opus users. From June 15, 2026: Chat + Agent SDK credit pools."},{"provider":"OpenAI","tier":"Plus","price_usd":20,"limits":"~80 GPT-5.4 msg/3h","models":"GPT-5.4 (limited) · GPT-5.4 Mini · o-series","features":"ChatGPT · GPTs · Codex CLI 30-150 tasks/5h"},{"provider":"OpenAI","tier":"Pro 5×","price_usd":100,"limits":"5× Plus quotas (new tier April 2026)","models":"GPT-5.4 Thinking unlimited","features":"Released as Anthropic Max competitor"},{"provider":"OpenAI","tier":"Pro","price_usd":200,"limits":"Effectively unlimited","models":"All including o3-pro","features":"Original premium tier"},{"provider":"Google","tier":"Gemini Advanced","price_usd":20,"limits":"Generous, soft caps","models":"Gemini 3.1 Pro · Gemini 3.6 Flash · Gemini 3.5 Flash-Lite","features":"Workspace integration · 1M context"},{"provider":"Google","tier":"Ultra","price_usd":100,"limits":"Higher quotas + Veo video","models":"All Gemini + research preview","features":"Veo 3 video · Project Mariner"},{"provider":"Mistral","tier":"Le Chat Pro","price_usd":22,"limits":"Generous","models":"Mistral Large 3 · Codestral","features":"EU jurisdiction"},{"provider":"xAI","tier":"SuperGrok","price_usd":30,"limits":"Generous","models":"Grok 4 · Grok 4 Heavy","features":"X integration"},{"provider":"DeepSeek","tier":"API only","price_usd":null,"limits":"Pay per token","models":"V4-Pro · V4-Flash","features":"Cheapest frontier API ($0.435/$0.87 permanent since May 22, 2026)"}],"agent_policies":[{"provider":"Anthropic","subscription_automated":"Prohibited","enforcement":"Active (OAuth blocked Apr 4)","first_party_exception":"`claude -p` pipe mode and Claude Code itself","api_required_for_automation":true,"cite":[35]},{"provider":"OpenAI","subscription_automated":"Prohibited (ToS)","enforcement":"Currently tolerated","first_party_exception":"Codex CLI uses subscription quota","api_required_for_automation":false},{"provider":"Google","subscription_automated":"Allowed in CLI","enforcement":"—","first_party_exception":"Gemini CLI free tier 1000 req/day","api_required_for_automation":false},{"provider":"Mistral","subscription_automated":"Allowed","enforcement":"—","first_party_exception":"—","api_required_for_automation":false}],"harnesses":[{"id":"claude-code","name":"Claude Code","vendor":"Anthropic","license":"proprietary","category":"CLI + IDE","stars":49000,"providers":["Anthropic"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":true,"computer_use":true,"lsp":true,"git":true,"memory":"CLAUDE.md","sandbox":"local","swe_pro":46,"pricing":"Subscription Pro/Max","sweet_spot":"SubagentStop hooks, /plugin list, requiredMinimumVersion managed setting, MCP fixes, Agent Teams.","stumbles":"Anthropic-only. Removed from standard Pro tier April 2026 — push to Max.","cite":[23,1]},{"id":"opencode","name":"OpenCode","vendor":"Anomaly","license":"MIT","category":"CLI + ACP","stars":188119,"providers":["75+: Anthropic*, OpenAI, Google, Mistral, Kimi, GLM, Ollama, LM Studio"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.18.4 (Jul 20, 2026): adaptive Kimi thinking controls, provider-defined reasoning options, restored Azure endpoints, and a rewritten desktop prompt input.","stumbles":"Anthropic OAuth blocked April 4 — must use API key.","cite":[24,35,72]},{"id":"codex-cli","name":"Codex CLI","vendor":"OpenAI","license":"Apache 2.0","category":"CLI + macOS app","stars":100232,"providers":["OpenAI"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"cloud","swe_pro":56.8,"pricing":"Plus $20 / Pro $200","sweet_spot":"v0.144.6 (Jul 18, 2026): refreshed GPT-5.6 Sol/Terra/Luna bundled instructions and corrected Codex context-window metadata to 272K.","stumbles":"OpenAI-only. No MCP, no hooks. Tightly coupled to apply_patch tool.","cite":[25,36,70]},{"id":"gemini-cli","name":"Gemini CLI","vendor":"Google","license":"Apache 2.0","category":"CLI","stars":106096,"providers":["Google"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"GEMINI.md","sandbox":"local","swe_pro":null,"pricing":"Free 1000 req/day","sweet_spot":"v0.51.0 (Jul 16, 2026): hardened sensitive-path and symlink handling, read-only macOS sandbox git config, and modern-model escape-sequence fixes.","stumbles":"Sunsetting to Antigravity CLI for free tier on June 18, 2026; paid Gemini/Enterprise keys retain access.","cite":[26,71]},{"id":"aider","name":"Aider","vendor":"paul-gauthier","license":"Apache 2.0","category":"CLI","stars":32000,"providers":["Anthropic*","OpenAI","Google","Ollama","100+"],"mcp":false,"skills":false,"hooks":false,"subagents":false,"voice":true,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"CONVENTIONS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"Mature, lightweight, voice-input native. Pair-programmer mode. Active, last commit Mar 2026. 44k stars.","stumbles":"No MCP, no hooks. Less ambitious than newer harnesses.","cite":[27]},{"id":"cline","name":"Cline","vendor":"cline-bot","license":"Apache 2.0","category":"VS Code extension","stars":64886,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".clinerules","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v4.0.10 (Jul 20, 2026): current release adds telemetry for consecutive-mistake-limit events; multi-editor and CLI surfaces remain available.","stumbles":"Anthropic OAuth blocked. Can be expensive on long sessions.","cite":[28,35,73]},{"id":"roo-code","name":"Roo Code","vendor":"RooVetGit","license":"Apache 2.0","category":"VS Code extension","stars":18000,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".roo","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v3.53.0 (Apr 23, 2026). Power-user Cline fork, model-agnostic, BYOK. Recent: GPT-5.5 via Codex, Opus 4.7 on Vertex, checkpoint nav.","stumbles":"Same OAuth situation. Configuration complexity.","cite":[29,35]},{"id":"cursor","name":"Cursor","vendor":"Anysphere","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".cursorrules","sandbox":"local","swe_pro":null,"pricing":"Hobby free · Pro $20 · Pro+ $60 · Ultra $200 · Teams $40/user","sweet_spot":"Cursor 3.5 (May 20, 2026): Cloud Agents (isolated VMs, multi-repo), Composer 2.5, Agents Window, parallel subagents.","stumbles":"Closed source. Lock-in. Subscription required for serious use.","cite":[31]},{"id":"windsurf","name":"Windsurf / Devin Desktop","vendor":"Cognition","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"windsurfrules","sandbox":"local","swe_pro":null,"pricing":"Pro $20 · Max $200","sweet_spot":"Renamed to Devin Desktop, Agent Command Center kanban, embedded Devin cloud agent, SWE-1.6 model, multi-agent + worktrees.","stumbles":"Pro $20 (was $15); smaller ecosystem than Cursor.","cite":[32]},{"id":"goose","name":"Goose","vendor":"Block","license":"Apache 2.0","category":"Desktop + CLI","stars":51387,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":false,"git":true,"memory":".goosehints","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.43.0 (Jul 14, 2026): per-message token/cost/TTFT/tok-s metrics, ACP reconnection, GPT-5.6 support, dynamic Ollama Cloud discovery, and expanded providers.","stumbles":"Anthropic OAuth blocked. Less mindshare than OpenCode.","cite":[30,35,74]},{"id":"omo","name":"OMO (Multi-model orchestrator)","vendor":"community","license":"MIT","category":"Multi-agent orchestrator","stars":54000,"providers":["multi-provider"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free","sweet_spot":"Rebrand from oh-my-opencode; multi-model orchestration. Wraps Claude Code, OpenCode, Codex, Kimi K2, DeepSeek V4, Gemini CLI.","stumbles":"Niche. Steep learning curve."},{"id":"hermes","name":"Hermes Agent","vendor":"Nous Research","license":"MIT","category":"Self-hosted multi-platform agent","stars":null,"providers":["OpenRouter-style multi-model"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":true,"remote":true,"computer_use":true,"lsp":false,"git":true,"memory":"persistent memory + auto-gen skills","sandbox":"local/docker/ssh/singularity/modal","swe_pro":null,"pricing":"Free (self-hosted)","sweet_spot":"v0.16.0. Persistent memory + auto-generated skills — learns your projects. Bridges Telegram/Discord/Slack/WhatsApp/Signal/Email/CLI. Natural-language cron for unattended runs. Parallel isolated subagents. Web search, browser automation, vision, image-gen, TTS.","stumbles":"Not coding-specialised — general autonomous agent. Setup requires self-host script. No first-party SWE benchmark.","cite":[]},{"id":"github-copilot-cli","name":"GitHub Copilot CLI","vendor":"GitHub","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["GitHub Models"],"mcp":false,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":true,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Bundled with Copilot Business/Enterprise","sweet_spot":"GA Feb 25, 2026. Specialized sub-agents (Explore, Task, Code Review, Plan), background delegation, autopilot.","stumbles":"Bundled-only — no standalone tier. GitHub-centric."},{"id":"amp-cli","name":"Amp CLI","vendor":"Sourcegraph","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["Multi"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Subscription (Sourcegraph)","sweet_spot":"Spun out as standalone company 2026. Runs as sidebar agent inside Zed via Terminal Threads.","stumbles":"Early standalone phase."},{"id":"zed","name":"Zed","vendor":"Zed Industries","license":"proprietary","category":"IDE","stars":null,"providers":["15 LLM providers + MCP"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Personal free (2k predictions) · Pro $10 · Business $30/seat","sweet_spot":"Rust-native editor with first-class agent panel + ACP host. Terminal Threads run Claude Code/Amp inline.","stumbles":"Editor first; agent layer still maturing."},{"id":"continue-dev","name":"Continue","vendor":"Continue","license":"Apache 2.0","category":"IDE extension + CLI","stars":null,"providers":["multi-provider"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".continue","sandbox":"local","swe_pro":null,"pricing":"Solo $0 · Team/Company ~$10/dev/mo","sweet_spot":"Agent mode plan+execute. Continuous AI, Mission Control, shared PR/ticket workflows.","stumbles":"Newer agent features still stabilizing across providers."}],"self_hosting":{"hardware_options":[{"id":"vast-2xa6000","name":"Vast.ai 2× A6000","vram_gb":96,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Reference cloud-GPU setup for this dashboard. No upfront capex. Docker templates, SSH/Cloudflare Zero Trust. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-3xa6000","name":"Vast.ai 3× A6000","vram_gb":144,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Headroom for Qwen 3 235B-A22B Q6_K. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-4xa6000","name":"Vast.ai 4× A6000","vram_gb":192,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Llama 3.1 405B Q4 territory. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"mbp-14-m4pro-64","name":"MacBook Pro 14\" M4 Pro 64GB","vram_gb":64,"cost_label":"~$3,200 capex","cost_per_hour":null,"type":"local","notes":"MoE sweet spot. 273 GB/s memory bandwidth."},{"id":"mbp-16-m5max-128","name":"MacBook Pro 16\" M5 Max 128GB","vram_gb":128,"cost_label":"~$5,500 capex","cost_per_hour":null,"type":"local","notes":"Best portable inference. ~545 GB/s bandwidth."},{"id":"mba-15-32","name":"MacBook Air 15\" 32GB","vram_gb":32,"cost_label":"~$1,900 capex","cost_per_hour":null,"type":"local","notes":"Hard 32GB ceiling. Limited to ~14B dense or 26B MoE Q4."},{"id":"contabo-2xl40s","name":"Contabo 2 x L40S","provider":"Contabo","vram_gb":96,"cost_label":"€1,502/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"EU-oriented fixed monthly configuration; 64 vCPU, 213 GB RAM, 3.5 TB storage, and 15 TB bandwidth are listed for the 2-GPU tier."},{"id":"contabo-1xh200","name":"Contabo 1 x H200 NVL","provider":"Contabo","vram_gb":141,"cost_label":"€2,149/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"Fixed monthly large-memory option. Confirm location and availability before treating it as a sovereignty or latency fit."},{"id":"infomaniak-1xl40s","name":"Infomaniak 1 x L40S","provider":"Infomaniak","vram_gb":48,"cost_label":"Live calculator / availability validation","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"calculator_or_request","price_checked":"2026-07-19","source":"https://www.infomaniak.com/en/hosting/public-cloud/prices","notes":"Swiss OpenStack option with dedicated GPU access and usage billing. Public pages list L40S availability but do not expose a stable crawlable SKU price; validate availability for the selected region."},{"id":"hyperstack-1xh200","name":"Hyperstack 1 x H200 SXM","provider":"Hyperstack","vram_gb":141,"cost_label":"$3.50/hr on demand","cost_per_hour":3.5,"type":"cloud","show_in_fit":false,"price_status":"published_on_demand","price_checked":"2026-07-19","source":"https://www.hyperstack.cloud/","notes":"Minute-accurate on-demand billing; reservation pricing starts at $2.45/hr. Validate region, storage, and availability."}],"models":[{"id":"gemma-4-26b-moe","name":"Gemma 4 26B-A4B MoE","params_total":26,"params_active":4,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":35,"livecodebench":77,"aime":88,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-3xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-4xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":22,"tok_per_sec":75},"mbp-16-m5max-128":{"quant":"BF16","vram_used":52,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"MoE — only 4B active per token. Faster than dense models 5× its size."},{"id":"gemma-4-31b-dense","name":"Gemma 4 31B Dense","params_total":31,"params_active":31,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":38,"livecodebench":80,"aime":89,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-3xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-4xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":19,"tok_per_sec":28},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":26,"tok_per_sec":38},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Highest quality open Gemma. Slower per-token (full 31B active)."},{"id":"qwen-3.6-plus","name":"Qwen 3.6 Plus","params_total":397,"params_active":17,"released":"2026-04-11","license":"Apache 2.0","jurisdiction":"China","swe_pro":50,"livecodebench":71.4,"aime":87,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":38},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":35},"vast-4xa6000":{"quant":"Q8_0","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1M context. Top open agentic coder. April 11 release.","swe_verified":68.2},{"id":"llama-4-maverick","name":"Llama 4 Maverick","params_total":400,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":42,"livecodebench":70,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":28},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":25},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"400B / 17B active MoE. 1M context. Strong MMLU-Pro (80.5%) but coding behind Chinese labs.","swe_verified":72},{"id":"minimax-m2.5-open","name":"MiniMax M2.5 (open weights)","params_total":456,"params_active":46,"released":"2026-01-20","license":"open-weight","jurisdiction":"China","swe_pro":40,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Eclipsed by DeepSeek V4 and Qwen 3.6 Plus. China jurisdiction. Maintained for niche workloads."},{"id":"deepseek-v4-pro-open","name":"DeepSeek V4-Pro (open weights)","params_total":1600,"params_active":49,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":55.4,"livecodebench":93.5,"aime":90,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q2_K","vram_used":188,"tok_per_sec":14},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1.6T total / 49B active MoE. MIT license. Strongest open coder. Requires very large multi-GPU deployment."},{"id":"llama-4-scout","name":"Llama 4 Scout","params_total":109,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":36,"swe_verified":68,"livecodebench":66,"aime":78,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-3xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-4xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":32,"tok_per_sec":22},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":45,"tok_per_sec":32},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"10M token context (longest in any open model). 109B / 17B active MoE."},{"id":"kimi-k2.6","name":"Kimi K2.6","params_total":235,"params_active":21,"released":"2026-05-01","license":"Modified MIT","jurisdiction":"China","swe_pro":47,"swe_verified":75,"livecodebench":78,"aime":84,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":36},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":32},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Best open-weight for sub-agent fan-out. Built for harness-driven parallel pipelines. Chinese-trained."},{"id":"glm-5.1","name":"GLM 5.1","params_total":358,"params_active":32,"released":"2026-04-22","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":76,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT license — rare among open frontier models besides DeepSeek. Strong for enterprise fine-tuning."},{"id":"mistral-small-4","name":"Mistral Small 4","params_total":24,"params_active":24,"released":"2026-04-18","license":"Apache 2.0","jurisdiction":"EU","swe_pro":28,"swe_verified":60,"livecodebench":58,"aime":68,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-3xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-4xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":48,"tok_per_sec":38},"mbp-16-m5max-128":{"quant":"BF16","vram_used":48,"tok_per_sec":55},"mba-15-32":{"quant":"Q4_K_M","vram_used":14,"tok_per_sec":30}},"notes":"6.5B effective parameters. EU jurisdiction. Best on-device option."},{"id":"nemotron-3-ultra-550b-a55b-moe","name":"Nvidia Nemotron 3 Ultra","params_total":550,"params_active":55,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q3_K_M","vram_used":180,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Hybrid Mamba-Transformer. 1M context. NVIDIA Open Model License. Computex June 4 launch."},{"id":"nemotron-3-nano-30b-a3b","name":"Nvidia Nemotron 3 Nano 30B-A3B","params_total":30,"params_active":3,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":30,"swe_verified":70,"livecodebench":68,"aime":76,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":24,"tok_per_sec":70},"mbp-16-m5max-128":{"quant":"BF16","vram_used":60,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":17,"tok_per_sec":50}},"notes":"Hybrid Mamba-Transformer Nano variant. Open weights."},{"id":"nemotron-nano-9b-v2","name":"Nvidia Nemotron Nano 9B v2","params_total":9,"params_active":9,"released":"2026-04-12","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":22,"swe_verified":60,"livecodebench":55,"aime":65,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-3xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-4xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":18,"tok_per_sec":90},"mbp-16-m5max-128":{"quant":"BF16","vram_used":18,"tok_per_sec":120},"mba-15-32":{"quant":"Q4_K_M","vram_used":6,"tok_per_sec":65}},"notes":"Dense 9B. Strong instruction following at edge sizes."},{"id":"kimi-k2.6-1t","name":"Kimi K2.6 (1T MoE)","params_total":1000,"params_active":32,"released":"2026-04-20","license":"Modified MIT","jurisdiction":"China","swe_pro":58.6,"swe_verified":80.2,"livecodebench":82,"aime":88,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Top open intelligence. 80.2% SWE-Verified, 58.6% SWE-Pro, AAII 54. 262K context. Modified MIT."},{"id":"glm-4.6","name":"GLM 4.6","params_total":355,"params_active":32,"released":"2025-09-15","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":75,"aime":80,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":28},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":24},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 200K context. Z.ai release."},{"id":"qwen3.6-35b-a3b","name":"Qwen 3.6 35B-A3B","params_total":35,"params_active":3,"released":"2026-04-16","license":"Apache 2.0","jurisdiction":"China","swe_pro":38,"swe_verified":74,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":22,"tok_per_sec":80},"mbp-16-m5max-128":{"quant":"BF16","vram_used":70,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"Apache 2.0 MoE — strong $/quality for fast bulk inference."},{"id":"cohere-command-a-plus","name":"Cohere Command A+","params_total":218,"params_active":28,"released":"2026-05-22","license":"CC-BY-NC + commercial","jurisdiction":"Canada","swe_pro":42,"swe_verified":75,"livecodebench":70,"aime":78,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":30},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":26},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"First open-weights Cohere release in &gt;1 year. Sparse-MoE multimodal."},{"id":"mistral-large-3-open","name":"Mistral Large 3 (open weights)","params_total":675,"params_active":41,"released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q3_K_M","vram_used":140,"tok_per_sec":22},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":180,"tok_per_sec":20},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"675B/41B active MoE. Apache 2.0. EU jurisdiction."},{"id":"deepseek-v4-flash-open","name":"DeepSeek V4-Flash (open weights)","params_total":284,"params_active":13,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":45},"vast-3xa6000":{"quant":"Q6_K","vram_used":115,"tok_per_sec":40},"vast-4xa6000":{"quant":"Q8_0","vram_used":150,"tok_per_sec":34},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":20},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 1M context, 384K max output. Cheap large-context open option."}],"frameworks":[{"name":"llama.cpp","best_for":"Mac (MLX), broad GGUF support","notes":"Best Apple Silicon performance via Metal."},{"name":"vLLM","best_for":"Multi-GPU servers, throughput","notes":"Production serving. Tensor parallelism."},{"name":"Ollama","best_for":"Easiest setup, dev workflow","notes":"Wrapper around llama.cpp. One-line model pull."},{"name":"MLX","best_for":"Apple Silicon native","notes":"Apple's framework. Best M-series perf."}]},"strategy":{"current_recommendation":{"label":"Risk-tiered GPT-5.6 + Fable stack","monthly_usd":null,"components":["GPT-5.6 Luna low — high-volume subagents and routine transformations","Gemini 3.5 Flash-Lite minimal — throughput-first extraction, parsing and parallel subagents","Gemini 3.6 Flash medium — Google-first coding, computer-use and multimodal agent loops","GPT-5.6 Terra medium — default engineering, review and documentation","GPT-5.6 Sol high or Claude Fable 5 high — escalation for hard, long-horizon tasks","Qwen 3.6 35B-A3B or DeepSeek V4-Flash — local route for privacy-sensitive bulk work"],"rationale":"Route by task risk instead of one subscription. Terra and Luna now retain strong coding-agent quality at substantially lower list price; Fable remains the long-horizon option only where its 30-day retention requirement is acceptable. Monthly cost is workload-dependent and must be computed from measured token volume."},"alternatives":[{"label":"Dual subscription (Claude Max + OpenAI Pro)","monthly_usd":230,"rationale":"Adds GPT-5.4 SWE-bench Pro lead and Codex CLI cloud sandbox. Worth $100/mo only if you frequently hit hard issues where Sonnet 4.6 plateaus.","verdict":"Defer until you have measured Sonnet plateau frequency for 30 days."},{"label":"API-only (no subscriptions)","monthly_usd":200,"rationale":"Pure pay-per-use. Maximum flexibility. Loses the subscription cache advantage — same workload costs 1.5–5× more for power users.","verdict":"Worse economics for your usage volume. Skip."},{"label":"Self-hosted maximalist","monthly_usd":80,"rationale":"Vast.ai 24/7 with Gemma 4 + Qwen 3 + occasional API top-up for frontier-only tasks.","verdict":"Cheapest if quality plateau at ~88% of Opus is acceptable. Operational overhead is real."}],"routing":[{"tier":"Bulk (70%)","use_for":"Classification, simple edits, triage, log parsing","preferred":"Gemini 3.5 Flash-Lite minimal · GPT-5.6 Luna low · Qwen 3.6 35B-A3B when local","cost_label":"Gemini $0.30/$2.50; Luna $0.20/$1.20 Mtok before cache, batch and reasoning tokens"},{"tier":"Mid (25%)","use_for":"Multi-file edits, code review, refactors, docs","preferred":"GPT-5.6 Terra medium · Gemini 3.6 Flash medium for Google-first or multimodal work","cost_label":"Terra $2/$12; Gemini $1.50/$7.50 Mtok before cache, batch and reasoning tokens"},{"tier":"Premium (5%)","use_for":"Architecture, hard debugging, long-context refactors","preferred":"GPT-5.6 Sol high · Claude Fable 5 high for long-horizon autonomy","cost_label":"Sol $5/$30; Fable $10/$50 Mtok before reasoning tokens"}],"open_questions":["Will the Nvidia Nemotron Coalition (Mistral, Cursor, Black Forest Labs, Thinking Machines) actually deliver a frontier-open consortium, or fragment within a quarter?","How does the June 15 Anthropic billing restructure (Chat pool + Agent SDK credit pool) change Max economics for agentic workloads?","Does Gemini 3.6 Flash’s lower output price and stronger agentic performance displace 3.1 Pro for most coding workflows before Gemini 3.5 Pro arrives?","DeepSeek V4-Pro permanent pricing ($0.435/$0.87) — does it force closed providers (OpenAI, Anthropic) to cut headline prices?","SubQ 1M-Preview's subquadratic architecture — does the next year see other commercial non-transformer LLMs?"],"task_fit":{"description":"For each task type, one recommended pick per provider drawn from the full roster. Click any cell to see runners-up that the daily briefing also considered. Burn × is derived from the matrix and recomputes when you change the reference dropdown.","rows":[{"task":"Tiny edits / boilerplate","description":"Single-file syntax fixes, regex replacements, formatting, docstrings, import sorting","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fixed-rate, fastest paid Anthropic option"},"openai":{"model_id":"gpt-5.5","effort":"minimal","rationale":"Cheapest reasoning setting on flagship"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Fastest Google route for cheap, high-volume pattern work"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Effectively $0/tok on existing Vast.ai infra"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ minimal","gpt-5.4-nano @ minimal"],"anthropic":["sonnet-4.6 @ low"],"google":["gemini-3.1-pro @ off"],"self_hosted":["llama-3.1-70b @ Q6_K"]}},{"task":"Normal coding task","description":"Single-function implementation, straightforward bug fixes, simple feature additions","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best $/quality on Anthropic side; cache reads included in Max 5×"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-specialised, ~⅓ burn of GPT-5.5 for near-identical SWE-Pro"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"58.7% SWE-Pro and stronger agentic coding than 3.5 Flash at lower output price"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Highest-quality open Gemma at full BF16 in 96GB"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium","gpt-5.5 @ low"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ high"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Multi-file implementation","description":"Feature spanning 3-10 files, requires understanding cross-file context","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Best style/intent understanding for code spanning files"},"openai":{"model_id":"gpt-5.5","effort":"medium","rationale":"Strong on planning; medium gives consistent multi-file edits"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Improved multi-step coding loops, fewer unwanted edits and 1M context"},"self_hosted":{"model_id":"qwen-3.6-plus","effort":null,"rationale":"Largest open MoE; strong reasoning across files"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ high","gpt-5.4 @ high"],"anthropic":["opus-4.7 @ high"],"google":["gemini-3.1-pro @ high"],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Hard debugging / architecture","description":"Race conditions, performance bottlenecks, design decisions with long-term implications","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Default Claude Code effort for Opus; best long-context retention"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"SWE-Pro leader; high effort balances depth and burn"},"google":{"model_id":"gemini-3.1-pro","effort":"high","rationale":"Preview Pro remains the deepest Google reasoning route; compare 3.6 Flash before paying the premium"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open-weight models trail frontier ~10-15pts here; not yet ready for hardest cases"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh","gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ high (lower burn)"],"google":[],"self_hosted":["qwen-3-235b-a22b — viable for some cases"]}},{"task":"Very hard autonomous repo task","description":"Overnight runs, agentic loops, tasks you don't supervise turn-by-turn","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"max","rationale":"Autonomy earns max effort's premium; no human in the loop to correct"},"openai":{"model_id":"gpt-5.5","effort":"xhigh","rationale":"Highest available reasoning budget on OpenAI side"},"google":{"model_id":null,"effort":null,"rationale":"Gemini 3.6 Flash improves agent loops but still lacks a max-equivalent tier for the hardest unsupervised runs"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Not recommended — frontier models earn their cost on the hardest tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ xhigh — better but $100/mo gating"],"anthropic":["opus-4.7 @ xhigh — if max budget is too steep"],"google":[],"self_hosted":[]}},{"task":"Code review / PR comments","description":"Reviewing diffs, suggesting improvements, finding issues without writing code yourself","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Critical thinking at moderate burn; doesn't need Opus depth"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-tuned reviewer at ⅓ burn of GPT-5.5"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Better instruction following and fewer unwanted code edits than 3.5 Flash"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Critique tasks don't need frontier; quality plateau acceptable"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ medium"],"self_hosted":["qwen-3-235b-a22b"]}},{"task":"Long-context refactor (50K+ tokens)","description":"Cross-cutting changes across a large codebase, where retention is the bottleneck","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"high","rationale":"1M context with retention that actually works; Opus's sweet spot"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"400K API context; degrades faster than Opus past 200K"},"google":{"model_id":"gemini-3.6-flash","effort":"high","rationale":"Google reports a 54% 1M-context MRCR score, roughly double 3.5 Flash"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open models cap around 200K useful context; not yet competitive"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh — if context fits in 256K"],"anthropic":["opus-4.7 @ xhigh"],"google":["gemini-3.1-pro @ high"],"self_hosted":[]}},{"task":"Subagent / delegated subtask","description":"Spawned by a main agent for focused work; quality threshold lower than user-facing","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fast, cheap, no effort knob to manage"},"openai":{"model_id":"gpt-5.4-mini","effort":"medium","rationale":"Best subagent — 94% of GPT-5.4 coding at 6× less"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Purpose-built for high-volume subagents at roughly 490 output tok/s"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Zero per-token cost for high-volume subagent loops"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ low","gpt-5.4-nano @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["llama-3.1-70b @ Q4_K_M"]}},{"task":"Documentation / explanation","description":"READMEs, code comments, tutorials, explaining existing code in prose","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best prose style on Anthropic side"},"openai":{"model_id":"gpt-5.4","effort":"medium","rationale":"Reasoning helps less here; medium-effort GPT-5.4 wins on $/quality"},"google":{"model_id":"gemini-3.6-flash","effort":"low","rationale":"Lower output price than 3.5 Flash with improved knowledge-work and document analysis"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Open models are competitive for prose tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ low"],"anthropic":["haiku-4.5"],"google":["gemini-3.5-flash-lite @ low"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Test scaffolding / fixtures","description":"Boilerplate test files, fixture generation, mocking, parameterised test suites","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Pattern-following work; doesn't need reasoning depth"},"openai":{"model_id":"gpt-5.4-mini","effort":"low","rationale":"Pattern-heavy; minimal reasoning, fast turnaround"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Lowest-latency current Google model for repetitive structured generation"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Bulk test scaffolding is exactly where self-hosted earns out"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Cybersecurity / vulnerability research","description":"Penetration testing, red-teaming, vulnerability discovery — tasks where cyber-specific safeguards apply","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Most capable Anthropic model; cyber tasks require Cyber Verification Program"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"Flagship reasoning; vetted security access via OpenAI"},"google":{"model_id":null,"effort":null,"rationale":"No cyber-specialized Gemini variant"},"self_hosted":{"model_id":"deepseek-v4-pro","effort":null,"rationale":"Open-weight without safety filtering; deploy in isolated environment"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ xhigh"],"google":[],"self_hosted":["qwen-3.6-plus","kimi-k2.6"]}}]}},"quota_burn_matrix":{"baseline_label":"selected reference model at medium effort = 1.00×","baseline_model_id":"fable-5","baseline_effort":"medium","methodology":"Burn ratios are displayed as multiples of the selected reference model at medium effort. Underlying values stored as absolute units anchored to gpt-5.5 medium = 1.00. Per-model ratios from API list pricing (firm, ±5%). Per-effort multipliers grounded in: Anthropic's published thinking-token budgets (low=skip, medium=~1k, high=5k, xhigh=10k, max=20k tokens); ArtificialAnalysis's measurement of Sonnet 4.6 max ≈ 3× Sonnet 4.5 on the Intelligence Index; nxcode.io / OpenAI guidance that xhigh ≈ 3-5× medium; ampcode's GPT-5.5 cost analysis. Per-effort multipliers are ±20%. NOTE: quality vs effort is not strictly monotonic per task — aggregate quality trends upward, but individual tasks can see high beat xhigh or medium beat high. OpenAI explicitly warns that 'high is not automatically better than medium'. Models that run in a separate preview quota bucket with no standard-pool burn are encoded as 0.00x and annotated until final token or credit rates are published.","stacking_multipliers":[{"name":"Fast mode (/fast on)","multiplier":2,"scope":"any cell"},{"name":"Cached input (long thread)","multiplier":0.6,"scope":"any cell"},{"name":"Plan mode (/plan)","multiplier":"uses high effort","scope":"overrides default"}],"openai_matrix":[{"model_id":"gpt-5.6-sol","minimal":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4,"note":"List-price blend equals the historical GPT-5.5 medium unit. Effort factors are planning estimates until measured token use is available."},{"model_id":"gpt-5.6-terra","minimal":0.18,"low":0.26,"medium":0.4,"high":0.68,"xhigh":1.08,"max":1.6,"note":"Two-fifths of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.6-luna","minimal":0.018,"low":0.026,"medium":0.04,"high":0.068,"xhigh":0.108,"max":0.16,"note":"Four percent of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.5","minimal":0.2,"low":0.5,"medium":1,"high":2,"xhigh":3.5},{"model_id":"gpt-5.5-pro","minimal":null,"low":null,"medium":6,"high":12,"xhigh":21},{"model_id":"gpt-5.4","minimal":0.1,"low":0.25,"medium":0.5,"high":1,"xhigh":1.75},{"model_id":"gpt-5.4-mini","minimal":0.011,"low":0.027,"medium":0.053,"high":0.107,"xhigh":0.187},{"model_id":"gpt-5.4-nano","minimal":0.003,"low":0.007,"medium":0.013,"high":null,"xhigh":null},{"model_id":"gpt-5.3-codex","minimal":0.07,"low":0.17,"medium":0.33,"high":0.67,"xhigh":1.17},{"model_id":"gpt-5.2","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null},{"model_id":"gpt-5.2-codex","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null}],"anthropic_matrix":[{"model_id":"fable-5","low":1.1,"medium":1.69,"high":2.87,"xhigh":4.56,"max":6.76,"note":"Always-on adaptive thinking. Values start from the $10/$50 list-price blend; effort multipliers are planning estimates rather than published token budgets."},{"model_id":"opus-4.8","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Current flagship. Same headline $5/$25. Fast Mode 3× cheaper ($10/$50) vs Opus 4.7."},{"model_id":"opus-4.7","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Legacy as of May 28. Thinking tokens: low=skip, medium=~1k, high=5k, xhigh=10k, max=20k. New tokenizer +35% tokens vs 4.6."},{"model_id":"sonnet-4.6","low":0.25,"medium":0.5,"high":1.25,"xhigh":null,"max":3.5,"note":"API default high. No xhigh tier. AA measured max ≈ 3× Sonnet 4.5 cost."},{"model_id":"haiku-4.5","low":null,"medium":0.17,"high":null,"xhigh":null,"max":null,"note":"No effort control — fixed-rate model."}],"google_matrix":[{"model_id":"gemini-3.1-pro","off":0.13,"low":0.2,"medium":0.3,"high":0.5},{"model_id":"gemini-3.5-flash","off":0.02,"low":0.03,"medium":0.05,"high":0.09},{"model_id":"gemini-3.6-flash","off":0.017,"low":0.025,"medium":0.042,"high":0.076},{"model_id":"gemini-3.5-flash-lite","off":0.005,"low":0.008,"medium":0.014,"high":0.024}],"non_reasoning_note":"Models without reasoning controls: Haiku 4.5 (Anthropic matrix above as fixed rate), GPT-5.5 Instant (ChatGPT default since May 5), Mistral Medium 3.5, MiniMax M2.7, Kimi K2.6, GLM 4.6, Qwen 3.6 variants, all Llama 4 variants, Grok 4.1 Fast, SubQ 1M-Preview — burn at fixed rate per model regardless of effort knob. DeepSeek V4-Flash-0731 now exposes thinking and non-thinking modes.","effort_quality_factors":{"minimal":0.65,"low":0.85,"medium":0.94,"high":0.98,"xhigh":1,"max":1.02,"off":0.85,"on":0.95},"quality_methodology":"Quality % = (model_swe_pro / reference_swe_pro × effort_quality_factor) × 100. Effort quality factors: minimal=0.65, low=0.85, medium=0.94, high=0.98, xhigh=1.00, max=1.02. Anchors: Anthropic's Hex measurement (low Opus 4.7 ≈ medium Opus 4.6 quality) anchors low ≈ 0.85; apiyi.com's report that max gains ~3pts over xhigh on hardest tasks anchors max=1.02; ampcode's GPT-5.5 analysis (medium captures 'most' of capability) anchors medium=0.94. IMPORTANT: these are AGGREGATE estimates. stet.sh found per-task reversals — high can beat xhigh on some tasks, medium can beat high. Treat as ±10% indicators, not precise measures.","unit_anchor":"gpt-5.5 medium","unit_anchor_note":"All burn values are stored as absolute units anchored to gpt-5.5 medium = 1.00. At render time, each cell is divided by the reference model's medium-effort cell (or its single datapoint for non-reasoning models) to produce the displayed ×-of-reference number. When the reference dropdown changes, all burn ratios recalculate against the new reference.","workload_presets":{"cold":{"label":"Cold (0% cache)","description":"One-off prompts, no context reuse. Worst-case burn.","cache_hit_rate":0},"mixed":{"label":"Mixed (40% cache)","description":"Interactive coding with some context reuse. Typical default.","cache_hit_rate":0.4},"warm":{"label":"Warm (70% cache)","description":"Sustained agentic session. AA's industry-standard 7:2:1 blend assumes this rate. danielvaughan.com's Codex CLI model uses 70%.","cache_hit_rate":0.7},"hot":{"label":"Hot (90% cache)","description":"Long warmer-pattern sessions with stable system prompts and reused context. vsits.co documented 99% achievable on Claude Code Max subscription with optimization.","cache_hit_rate":0.9}},"default_workload":"mixed","cache_methodology":"Cache discount = (cache_hit_price / input_price). Lower = better discount. Effective burn = (1 - cache_hit_rate) × raw_burn + cache_hit_rate × raw_burn × cache_discount. Per-provider cache discounts: Anthropic 10% (90% off), OpenAI 25% on flagship/10% on 5.4 family, DeepSeek V4-Pro 0.83% (99.2% off — most aggressive in industry), DeepSeek V4-Flash 2%, MiniMax 52%, Mistral &amp; Grok no published cache pricing (modeled as 100% = no discount), Google ~25% plus per-hour storage fee (not modeled). IMPORTANT CAVEAT on Anthropic subscriptions: docs say cache reads count at 10% rate, but GitHub issue anthropics/claude-code#24147 reports cache reads burning quota at full rate on Max subscriptions. Treat the displayed cache benefit as accurate for API usage; subscription quota behavior is contested. Cache writes (1.25× / 2× of input rate on Anthropic) are not modeled — amortize across many reads in steady-state."},"sources":[{"n":1,"category":"Anthropic","title":"Anthropic news","url":"https://www.anthropic.com/news"},{"n":2,"category":"Anthropic","title":"Anthropic Claude model overview","url":"https://docs.claude.com/en/docs/about-claude/models/overview"},{"n":3,"category":"Anthropic","title":"Anthropic pricing","url":"https://www.anthropic.com/pricing"},{"n":4,"category":"OpenAI","title":"OpenAI news","url":"https://openai.com/news/"},{"n":5,"category":"OpenAI","title":"OpenAI model catalog","url":"https://platform.openai.com/docs/models"},{"n":6,"category":"OpenAI","title":"OpenAI API pricing","url":"https://openai.com/api/pricing/"},{"n":7,"category":"Google","title":"Google DeepMind blog","url":"https://blog.google/technology/google-deepmind/"},{"n":8,"category":"Google","title":"Gemini API models","url":"https://ai.google.dev/gemini-api/docs/models"},{"n":9,"category":"Google","title":"Gemini API pricing","url":"https://ai.google.dev/pricing"},{"n":10,"category":"Mistral","title":"Mistral news","url":"https://mistral.ai/news/"},{"n":11,"category":"Mistral","title":"Mistral La Plateforme pricing","url":"https://mistral.ai/products/la-plateforme#pricing"},{"n":12,"category":"xAI","title":"xAI blog (Grok)","url":"https://x.ai/blog"},{"n":13,"category":"DeepSeek","title":"DeepSeek API docs","url":"https://api-docs.deepseek.com/"},{"n":14,"category":"MiniMax","title":"MiniMax news","url":"https://www.minimaxi.com/en/news"},{"n":15,"category":"Benchmarks","title":"SWE-bench (Verified)","url":"https://www.swebench.com/","note":"SWE-bench Verified is widely considered contaminated; treat with skepticism."},{"n":16,"category":"Benchmarks","title":"SWE-bench Pro leaderboard","url":"https://scale.com/leaderboard/swe_bench_pro","note":"Preferred trustworthy code-agent benchmark."},{"n":17,"category":"Benchmarks","title":"LiveCodeBench","url":"https://livecodebench.github.io/"},{"n":18,"category":"Benchmarks","title":"Artificial Analysis (cross-provider pricing + benchmarks)","url":"https://artificialanalysis.ai/"},{"n":19,"category":"Benchmarks","title":"LMArena (chatbot arena)","url":"https://lmarena.ai/"},{"n":20,"category":"Open-weight","title":"Hugging Face — trending models","url":"https://huggingface.co/models?sort=trending"},{"n":21,"category":"Open-weight","title":"Hugging Face blog","url":"https://huggingface.co/blog"},{"n":22,"category":"Open-weight","title":"r/LocalLLaMA (community signal)","url":"https://www.reddit.com/r/LocalLLaMA/"},{"n":23,"category":"Harnesses","title":"Claude Code (anthropics/claude-code)","url":"https://github.com/anthropics/claude-code"},{"n":24,"category":"Harnesses","title":"OpenCode (sst/opencode)","url":"https://github.com/sst/opencode"},{"n":25,"category":"Harnesses","title":"Codex CLI (openai/codex)","url":"https://github.com/openai/codex"},{"n":26,"category":"Harnesses","title":"Gemini CLI (google-gemini/gemini-cli)","url":"https://github.com/google-gemini/gemini-cli"},{"n":27,"category":"Harnesses","title":"Aider (Aider-AI/aider)","url":"https://github.com/Aider-AI/aider"},{"n":28,"category":"Harnesses","title":"Cline (cline/cline)","url":"https://github.com/cline/cline"},{"n":29,"category":"Harnesses","title":"Roo Code (RooVetGit/Roo-Code)","url":"https://github.com/RooVetGit/Roo-Code"},{"n":30,"category":"Harnesses","title":"Goose (block/goose)","url":"https://github.com/block/goose"},{"n":31,"category":"Harnesses","title":"Cursor changelog","url":"https://cursor.com/changelog"},{"n":32,"category":"Harnesses","title":"Windsurf changelog","url":"https://windsurf.com/changelog"},{"n":33,"category":"Hardware","title":"Vast.ai (spot GPU pricing — A6000, H100)","url":"https://vast.ai/"},{"n":34,"category":"Hardware","title":"Apple MacBook Pro (M-series)","url":"https://www.apple.com/shop/buy-mac/macbook-pro"},{"n":35,"category":"Policy","title":"Anthropic OAuth third-party restrictions (Apr 4, 2026)","url":"https://www.anthropic.com/news"},{"n":36,"category":"Policy","title":"OpenAI Codex token-based pricing migration (Apr 2, 2026)","url":"https://openai.com/news/"},{"n":37,"category":"Aggregators","title":"Hacker News (AI tags)","url":"https://news.ycombinator.com/"},{"n":38,"category":"Benchmarks","title":"AIME (math competition benchmark)","url":"https://aimeproblems.com/","note":"Used to anchor Reasoning axis ratings."},{"n":39,"category":"Benchmarks","title":"MMLU / MMLU-Pro (knowledge benchmark)","url":"https://github.com/hendrycks/test","note":"Used to anchor Knowledge axis ratings."},{"n":40,"category":"Benchmarks","title":"tau-bench (tool-use + agentic behaviour)","url":"https://github.com/sierra-research/tau-bench","note":"Used to anchor Agentic axis ratings."},{"n":41,"category":"Nvidia","title":"build.nvidia.com","url":"https://build.nvidia.com/"},{"n":42,"category":"Nvidia","title":"Hugging Face — Nvidia","url":"https://huggingface.co/nvidia"},{"n":43,"category":"Nvidia","title":"Nvidia blogs","url":"https://blogs.nvidia.com/"},{"n":44,"category":"Benchmarks","title":"SWE-bench Pro public leaderboard (Scale)","url":"https://scale.com/leaderboard/swe_bench_pro_public"},{"n":45,"category":"Benchmarks","title":"Scale Labs","url":"https://labs.scale.com/"},{"n":46,"category":"Benchmarks","title":"SWE-Rebench","url":"https://swe-rebench.com/"},{"n":47,"category":"Benchmarks","title":"Terminal-Bench 2.0","url":"https://tbench.ai/leaderboard"},{"n":48,"category":"Aggregators","title":"OpenRouter","url":"https://openrouter.ai/"},{"n":49,"category":"Aggregators","title":"LLM-Stats","url":"https://llm-stats.com/"},{"n":50,"category":"Benchmarks","title":"Artificial Analysis Intelligence Index","url":"https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index"},{"n":51,"category":"Harnesses","title":"Claude Code changelog","url":"https://code.claude.com/docs/en/changelog"},{"n":52,"category":"Harnesses","title":"Cursor pricing","url":"https://cursor.com/pricing"},{"n":53,"category":"Harnesses","title":"Windsurf changelog (Devin Desktop)","url":"https://windsurf.com/changelog"},{"n":54,"category":"Harnesses","title":"Gemini CLI changelogs","url":"https://geminicli.com/docs/changelogs"},{"n":55,"category":"Harnesses","title":"Codex CLI releases","url":"https://github.com/openai/codex/releases"},{"n":56,"category":"Google","title":"DeepMind models","url":"https://deepmind.google/models"},{"n":57,"category":"Mistral","title":"Mistral pricing","url":"https://mistral.ai/pricing"},{"n":58,"category":"Google","title":"Gemini API pricing (canonical)","url":"https://ai.google.dev/gemini-api/docs/pricing"},{"n":59,"category":"MiniMax","title":"Artificial Analysis — MiniMax M2.7","url":"https://artificialanalysis.ai/models/minimax-m2-7"},{"n":60,"category":"Moonshot","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3"},{"n":61,"category":"Moonshot","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models"},{"n":62,"category":"OpenAI","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/"},{"n":63,"category":"Hosting","title":"Runpod GPU cloud pricing","url":"https://www.runpod.io/pricing"},{"n":64,"category":"Hosting","title":"Runpod July 2024 GPU price changes","url":"https://www.runpod.io/blog/runpod-slashes-gpu-prices-more-power-less-cost-for-ai-builders"},{"n":65,"category":"Hosting","title":"Contabo GPU Cloud configurations and pricing","url":"https://contabo.com/en/gpu-cloud/"},{"n":66,"category":"Hosting","title":"Infomaniak Public Cloud pricing","url":"https://www.infomaniak.com/en/hosting/public-cloud/prices"},{"n":67,"category":"Hosting","title":"Infomaniak GPU flavor documentation","url":"https://docs.infomaniak.cloud/compute/instances/flavors/"},{"n":68,"category":"Hosting","title":"Hyperstack GPU cloud pricing","url":"https://www.hyperstack.cloud/"},{"n":69,"category":"Hosting","title":"Vast.ai marketplace pricing methodology","url":"https://docs.vast.ai/guides/instances/pricing"},{"n":70,"category":"Harnesses","title":"Codex CLI 0.144.6 release","url":"https://github.com/openai/codex/releases/tag/rust-v0.144.6"},{"n":71,"category":"Harnesses","title":"Gemini CLI 0.51.0 release","url":"https://github.com/google-gemini/gemini-cli/releases/tag/v0.51.0"},{"n":72,"category":"Harnesses","title":"OpenCode 1.18.4 release","url":"https://github.com/anomalyco/opencode/releases/tag/v1.18.4"},{"n":73,"category":"Harnesses","title":"Cline 4.0.10 release","url":"https://github.com/cline/cline/releases/tag/v4.0.10"},{"n":74,"category":"Harnesses","title":"Goose 1.43.0 release","url":"https://github.com/aaif-goose/goose/releases/tag/v1.43.0"},{"n":75,"category":"Moonshot","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing"},{"n":76,"category":"Swiss AI Initiative","title":"Apertus v1.1 4B Instruct model card","url":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"n":77,"category":"Google","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"n":78,"category":"Google","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"n":79,"category":"Google","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"n":80,"category":"Google","title":"Gemini API release notes","url":"https://ai.google.dev/gemini-api/docs/changelog"},{"n":81,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash"},{"n":82,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"},{"n":83,"category":"OpenAI","title":"GPT-5.6 launch and benchmark table","url":"https://openai.com/index/gpt-5-6/"},{"n":84,"category":"Google","title":"Gemini API Managed Agents update","url":"https://blog.google/innovation-and-ai/technology/developers-tools/expanding-managed-agents-gemini-api/"},{"n":85,"category":"Benchmarks","title":"Scale Labs SWE-bench Pro public leaderboard","url":"https://labs.scale.com/api/pdf/leaderboard/swe_bench_pro_public","note":"Public leaderboard values use their own model versions and evaluation setup; do not merge them mechanically with vendor launch tables."},{"n":86,"category":"Anthropic","title":"What's new in Claude Opus 5","url":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5","note":"Official launch specifications, pricing, availability, and migration behaviour."},{"n":87,"category":"Benchmarks","title":"Artificial Analysis — Claude Opus 5","url":"https://artificialanalysis.ai/models/claude-opus-5","note":"Independent effort-specific intelligence, latency, throughput, and price analysis."},{"n":88,"category":"Thinking Machines Lab","title":"Inkling model card","url":"https://thinkingmachines.ai/model-card/inkling/"},{"n":89,"category":"OpenAI","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra"},{"n":90,"category":"OpenAI","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna"},{"n":91,"category":"DeepSeek","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","note":"Vendor-reported benchmark results use DeepSeek Harness minimal mode at max effort; DSBench results are internal."},{"n":92,"category":"DeepSeek","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/"},{"n":93,"category":"OpenAI","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","note":"Official July 30 pricing, paid-subscription credit, and Sol API Fast mode details."},{"n":94,"category":"OpenAI","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","note":"Official August 6 ChatGPT product update; it does not change the API model contract."}],"provider_sources":{"Anthropic":[1,2,3,86,87],"OpenAI":[4,5,6,62,83,89,90,93,94],"Google":[7,8,9,58,77,78,79,80,81,82,84],"Mistral":[10,11],"xAI":[12],"DeepSeek":[13,91,92],"MiniMax":[14],"Meta":[20,21],"Alibaba":[20,21],"Moonshot":[20,21,60,61,75],"Zhipu":[20,21],"Nvidia":[41,42,43],"Cohere":[20,21],"Swiss AI Initiative":[76],"SubQ":[],"Thinking Machines Lab":[88]},"section_sources":{"models":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,60,61,62,76,77,78,79,80,81,82,86,87,89,90,91,92],"harnesses":[23,24,25,26,27,28,29,30,31,32,35,51,52,53,54,55,70,71,72,73,74],"self_hosting":[20,21,22,33,34,63,64,65,66,67,68,69,76],"strategy":[1,3,4,6,16,18,33,35,36]},"capabilities":{"schema_version":"1.0","axes":[{"key":"coding","label":"Coding","short":"Code","effort_sensitivity":0.5},{"key":"reasoning","label":"Reasoning &amp; Architecture","short":"R&amp;A","effort_sensitivity":0.6},{"key":"knowledge","label":"Knowledge &amp; Research","short":"K&amp;R","effort_sensitivity":0.05},{"key":"comms","label":"Communication &amp; Docs","short":"Comms","effort_sensitivity":0.2},{"key":"multimodal","label":"Multimodal","short":"MM","effort_sensitivity":0.05},{"key":"agentic","label":"Agentic","short":"Agent","effort_sensitivity":0.4}],"level_labels":{"coding":{"1":"Snippet","2":"Standard","3":"Cross-file","4":"Hard","5":"Frontier"},"reasoning":{"1":"Apply known","2":"Multi-step","3":"Cross-cutting","4":"Novel system","5":"Research-grade"},"knowledge":{"1":"Recall","2":"Contextual","3":"Cross-domain","4":"Frontier","5":"Original"},"comms":{"1":"Grammatical","2":"Structured","3":"Tutorial","4":"Editorial","5":"Publishable"},"multimodal":{"1":"Text only","2":"Image-in","3":"Image reasoning","4":"Video/audio","5":"Cross-modal gen"},"agentic":{"1":"Single-turn","2":"Tool calls","3":"Plan coherence","4":"Self-correcting","5":"Long-horizon"}},"focus_presets":{"balanced":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"coding-focused":{"coding":2,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"architecture-focused":{"coding":1,"reasoning":2,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"research-focused":{"coding":1,"reasoning":1,"knowledge":2,"comms":1,"multimodal":1,"agentic":1},"writing-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":2,"multimodal":1,"agentic":1},"multimodal-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":2,"agentic":1},"agentic-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":2}},"default_focus":"balanced","methodology":"1-5 levels per axis; Coding rated against SWE-Pro / SWE-Verified evidence; Reasoning against AIME / LiveCodeBench / public hard-reasoning evals; Multimodal against published modality support (image/audio/video in/out); Agentic rated holding harness constant at Claude-Code-class baseline (plan coherence, tool-call quality, self-correction, calibrated stopping, refusal hygiene); Knowledge from model-card claims + community-validated state-of-art awareness; Comms from prose-evaluation rounds + structural quality. Ratings carry citation; see src/dashboard-context.md for the full rubric.","effort_formula":"effective(axis, effort) = max(0, ceiling[axis] × (1 − effort_sensitivity[axis] × (1 − radar_effort_factor[effort]))). radar_effort_factor[medium] = 1.00 → primary polygon equals reference polygon shape at medium effort for the same model. Lower efforts shrink the polygon by sensitivity-weighted amounts; higher efforts grow it and may push vertices OUTSIDE the outer level-50 ring — the ring is a rubric marker, not a hard cap. Vertices that exceed 50 are drawn in accent-hot to signal 'boosted above medium-effort baseline'.","radar_effort_factors":{"minimal":0.5,"low":0.75,"medium":1,"high":1.08,"xhigh":1.15,"max":1.22},"scale":{"rubric_min":1,"rubric_max":50,"visual_max":60,"band_size":10,"boost_band":"51–60","interpretation":"capability_levels[axis] = the model's effective capability at MEDIUM effort. Briefings rate from 1 to 50. The radar's visual scale extends to 60 to accommodate the 'effort-boost band' (51–60) — values that fall there are formula-derived from extra reasoning effort, never directly rated. The outer rubric ring at 50 is visually emphasized as 'current frontier'.","bands":{"1-10":"Snippet","11-20":"Standard","21-30":"Cross-file","31-40":"Hard","41-50":"Frontier","51-60":"Effort-boost band (derived, not rated)"}}},"report_metrics":{"schema_version":"1.0","researched_at":"2026-08-01","currency":"USD","reference_options":["fable-5","gpt-5.6-sol","gpt-5.6-terra","gpt-5.6-luna","opus-4.8","gpt-5.5","gpt-5.5-pro","gpt-5.5-instant","kimi-k3","gemini-3.6-flash","gemini-3.5-flash-lite"],"default_reference":"fable-5","methodology":{"benchmark_scale":"Each benchmark is divided by benchmark.max_value and expressed on a 0-100 scale. The quality composite is the weighted arithmetic mean of available normalized benchmarks; it is shown only when coverage is at least 0.50.","quality_weights":{"aaii_v4_1":0.2,"coding_agent_index_v1_1":0.2,"swe_bench_pro":0.2,"deep_swe_v1_1":0.15,"terminal_bench_2_1":0.15,"agents_last_exam":0.1},"speed_score":"100 * sqrt(task_speed_index / max_task_speed_index). task_speed_index is 100 * reference_time / model_time. It is not output tokens per second and is comparable only within the cited evaluation setup.","cost_score":"10 + 90 * ln(max_blended_price / blended_price) / ln(max_blended_price / min_blended_price). Blended price is 0.30 * input_price + 0.70 * output_price per million tokens. Higher cost score is better.","scq_compound":"Geometric mean of quality_score, speed_score and cost_score. Geometric mean prevents one category from fully compensating for a weak category. Null when quality coverage is below 0.50.","capability_compound":"Weighted arithmetic mean of the six capability axes. Balanced uses weight 1 for each axis; a selected focus uses weight 2.5 for that axis and 1 for every other axis.","reference_quality":"Each model stores or derives a Fable-anchored quality value. Displayed quality = 100 * model_anchor / selected_reference_anchor, so the selected reference is always exactly 100. New-suite composites use only explicitly overlapping evaluations; legacy rows fall back to SWE-Bench Pro. Unknown evidence remains null, never zero.","burn":"Raw blended-price units are multiplied by an effort factor and cache factor, then divided by the selected reference model at medium effort. Cache factor = (1-hit_rate) + hit_rate*cache_read_ratio.","missing_data":"Keep unknown values null. Display insufficient comparable evidence rather than zero. Detailed benchmark cells remain version-specific and may be empty even when a reference-relative quality anchor exists from a separate documented comparison set."},"market_signals":[{"id":"frontier_quality","label":"Frontier quality index","explanation":"Best broad intelligence score published on Artificial Analysis Intelligence Index v4.1. Higher is better; the benchmark version must remain fixed across the series.","unit":"index","direction":"up","current":59.9,"history":[{"date":"2026-01-01","value":48.2},{"date":"2026-02-01","value":50.1},{"date":"2026-03-01","value":52.7},{"date":"2026-04-01","value":55.7},{"date":"2026-05-01","value":55.7},{"date":"2026-06-01","value":59.9},{"date":"2026-07-01","value":59.9}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"coding_quality","label":"Coding-agent frontier","explanation":"Best score on Artificial Analysis Coding Agent Index v1.1. Higher means better end-to-end coding-agent performance, not just code completion.","unit":"index","direction":"up","current":80,"history":[{"date":"2026-01-01","value":66.1},{"date":"2026-02-01","value":69.8},{"date":"2026-03-01","value":72.5},{"date":"2026-04-01","value":76.4},{"date":"2026-05-01","value":77.2},{"date":"2026-06-01","value":77.2},{"date":"2026-07-01","value":80}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"task_speed","label":"Coding task speed increase","explanation":"Fastest current end-to-end coding task-rate index, where Fable 5 is fixed at 100. A value of 320 means an estimated 3.2 times as many comparable tasks per unit time; it is not tokens per second.","unit":"Fable=100","direction":"up","current":320,"history":[{"date":"2026-01-01","value":100},{"date":"2026-02-01","value":112},{"date":"2026-03-01","value":126},{"date":"2026-04-01","value":145},{"date":"2026-05-01","value":180},{"date":"2026-06-01","value":180},{"date":"2026-07-01","value":320}],"source":"https://openai.com/index/gpt-5-6/","history_note":"Pre-July points are frozen planning estimates from prior report snapshots and should not be interpreted as one controlled longitudinal benchmark."},{"id":"blended_frontier_price","label":"Lowest frontier blended API price","explanation":"Lowest 30% input / 70% output list-price blend among models meeting the report's frontier quality floor. Lower is better; batch and caching are excluded.","unit":"$/Mtok","direction":"down","current":0.9,"history":[{"date":"2026-01-01","value":18},{"date":"2026-02-01","value":16.5},{"date":"2026-03-01","value":14},{"date":"2026-04-01","value":11.4},{"date":"2026-05-01","value":9},{"date":"2026-06-01","value":9},{"date":"2026-07-01","value":4.5},{"date":"2026-08-01","value":0.9}],"source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"open_weight_quality","label":"Open-weight quality vs selected reference","explanation":"Best open-weight model relative to the selected reference. Stored history is Fable-anchored and rebased in the browser whenever the reference changes. July uses Kimi K3's geometric mean across 14 overlapping launch-suite values; vendor-reported values require independent reproduction.","unit":"%","direction":"up","current":99.2,"history":[{"date":"2026-01-01","value":82},{"date":"2026-02-01","value":85},{"date":"2026-03-01","value":88},{"date":"2026-04-01","value":91},{"date":"2026-05-01","value":94},{"date":"2026-06-01","value":96},{"date":"2026-07-01","value":99.2}],"history_note":"Earlier points are frozen report estimates; July uses Moonshot’s K3 launch suite and is not a single controlled longitudinal benchmark.","source":"https://www.kimi.com/blog/kimi-k3"},{"id":"output_throughput","label":"Documented API output throughput","explanation":"Fastest provider-documented general API output rate in the tracked set. Unit is generated tokens per second; it is distinct from time to first token and end-to-end task speed.","unit":"tok/s","direction":"up","current":490,"history":[{"date":"2026-01-01","value":110},{"date":"2026-02-01","value":130},{"date":"2026-03-01","value":140},{"date":"2026-04-01","value":180},{"date":"2026-05-01","value":200},{"date":"2026-06-01","value":220},{"date":"2026-07-01","value":490}],"history_note":"Provider and independent figures use different serving and context conditions. The latest point is Artificial Analysis's Gemini 3.5 Flash-Lite measurement on Google's first-party API.","source":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"}],"benchmarks":[{"id":"aaii_v4_1","label":"AA Intelligence Index v4.1","max_value":59.9,"unit":"index","good_direction":"high"},{"id":"coding_agent_index_v1_1","label":"AA Coding Agent Index v1.1","max_value":80,"unit":"index","good_direction":"high"},{"id":"swe_bench_pro","label":"SWE-Bench Pro","max_value":80,"unit":"%","good_direction":"high"},{"id":"deep_swe_v1_1","label":"DeepSWE v1.1","max_value":72.7,"unit":"%","good_direction":"high"},{"id":"terminal_bench_2_1","label":"Terminal-Bench 2.1","max_value":91.9,"unit":"%","good_direction":"high","note":"Maximum is GPT-5.6 Sol Ultra; ordinary Sol scores 88.8."},{"id":"agents_last_exam","label":"Agents' Last Exam","max_value":52.7,"unit":"%","good_direction":"high"}],"model_metrics":[{"model_id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","benchmarks":{"aaii_v4_1":59.9,"coding_agent_index_v1_1":77.2,"swe_bench_pro":80,"deep_swe_v1_1":69.7,"terminal_bench_2_1":83.1,"agents_last_exam":40.5},"quality_score":94.9,"quality_coverage":1,"task_speed_index":100,"speed_score":50,"api_input":10,"api_cached_input":1,"api_output":50,"blended_price":38,"cost_score":10,"scq_compound":36.2,"speed_evidence":"Reference index; Anthropic labels comparative latency slower.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","benchmarks":{"aaii_v4_1":58.9,"coding_agent_index_v1_1":80,"swe_bench_pro":64.6,"deep_swe_v1_1":72.7,"terminal_bench_2_1":88.8,"agents_last_exam":52.7},"quality_score":95.4,"quality_coverage":1,"task_speed_index":256,"speed_score":80,"api_input":5,"api_cached_input":0.5,"api_output":30,"blended_price":22.5,"cost_score":22.6,"scq_compound":55.7,"speed_evidence":"OpenAI introduced API Fast mode on July 30, claiming up to 2.5× Standard speed at 2× price with no intelligence change; the base task-speed index remains a separate vendor comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","benchmarks":{"aaii_v4_1":55,"coding_agent_index_v1_1":77.4,"swe_bench_pro":63.4,"deep_swe_v1_1":69.6,"terminal_bench_2_1":87.4,"agents_last_exam":50.4},"quality_score":91.8,"quality_coverage":1,"task_speed_index":300,"speed_score":86.6,"api_input":2,"api_cached_input":0.2,"api_output":12,"blended_price":9,"cost_score":44.6,"scq_compound":70.8,"speed_evidence":"OpenAI reports roughly one-third of Fable 5 task time for the family coding comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","benchmarks":{"aaii_v4_1":51.2,"coding_agent_index_v1_1":74.6,"swe_bench_pro":62.7,"deep_swe_v1_1":67.2,"terminal_bench_2_1":84.7,"agents_last_exam":50.3},"quality_score":88.7,"quality_coverage":1,"task_speed_index":320,"speed_score":89.4,"api_input":0.2,"api_cached_input":0.02,"api_output":1.2,"blended_price":0.9,"cost_score":100,"scq_compound":92.6,"speed_evidence":"OpenAI calls Luna the fastest tier; 320 is a conservative planning index above the cited 300 family comparison and must not be treated as measured tok/s.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","benchmarks":{"deep_swe_v1_1":67.5,"terminal_bench_2_1":88.3},"quality_score":null,"quality_coverage":0.33,"quality_vs_fable":99.2,"task_speed_index":null,"speed_score":null,"api_input":3,"api_cached_input":0.3,"api_output":15,"blended_price":11.4,"cost_score":38.9,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time or output-throughput figure published for K3 at launch.","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"model_id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","benchmarks":{"aaii_v4_1":55.7,"coding_agent_index_v1_1":72.5,"swe_bench_pro":69.2,"deep_swe_v1_1":59,"terminal_bench_2_1":78.9,"agents_last_exam":45.2},"quality_score":87.7,"quality_coverage":1,"task_speed_index":180,"speed_score":67.1,"api_input":5,"api_cached_input":0.5,"api_output":25,"blended_price":19,"cost_score":26.7,"scq_compound":54,"speed_evidence":"Planning index based on moderate vendor latency and optional fast mode; validate in the local harness.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gemini-3.1-pro","name":"Gemini 3.1 Pro Preview","provider":"Google","benchmarks":{"aaii_v4_1":46.5,"coding_agent_index_v1_1":42.7,"swe_bench_pro":54.2,"deep_swe_v1_1":11.8,"terminal_bench_2_1":70.7,"agents_last_exam":32.1},"quality_score":59.8,"quality_coverage":1,"task_speed_index":null,"speed_score":null,"api_input":null,"api_cached_input":null,"api_output":null,"blended_price":null,"cost_score":null,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time measurement in the selected source set.","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro"},{"model_id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","benchmarks":{"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78},"quality_score":77.4,"quality_coverage":0.7,"quality_vs_fable":81.6,"task_speed_index":null,"speed_score":null,"api_input":1.5,"api_cached_input":0.15,"api_output":7.5,"blended_price":5.7,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 304 output tok/s at high thinking; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"model_id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","benchmarks":{"aaii_v4_1":36,"swe_bench_pro":54.2,"terminal_bench_2_1":54},"quality_score":62.5,"quality_coverage":0.55,"quality_vs_fable":65.9,"task_speed_index":null,"speed_score":null,"api_input":0.3,"api_cached_input":0.03,"api_output":2.5,"blended_price":1.84,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 490 output tok/s; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"}],"visualizations":{"bubble":{"x":"cost_score","y":"speed_score","size":"quality_vs_selected_reference","color":"provider","include_when":"quality vs reference, speed score and cost score are all non-null"},"heatmap":{"rows":"model","columns":["quality_vs_selected_reference","speed_score","cost_score"],"color_scale":"sequential_good","domain":[0,100]},"bcg":{"x":"capability_compound","y":"mean(quality_vs_selected_reference, speed_score, cost_score)","color":"provider","quadrants":"medians of currently visible models"}},"capability":{"axes":["coding","reasoning_architecture","knowledge_research","communication_docs","multimodal","agentic"],"focus_multiplier":2.5,"models":[{"model_id":"fable-5","scores":[98,99,96,96,90,99],"capability_compound":96.3,"scq_compound":36.2},{"model_id":"gpt-5.6-sol","scores":[98,97,95,95,96,98],"capability_compound":96.5,"scq_compound":55.7},{"model_id":"gpt-5.6-terra","scores":[95,92,91,92,90,95],"capability_compound":92.5,"scq_compound":70.8},{"model_id":"gpt-5.6-luna","scores":[92,88,86,89,86,91],"capability_compound":88.7,"scq_compound":92.6},{"model_id":"opus-4.8","scores":[92,94,94,95,82,94],"capability_compound":91.8,"scq_compound":54},{"model_id":"gemini-3.1-pro","scores":[78,84,94,84,98,72],"capability_compound":85,"scq_compound":null}],"provenance_note":"Capability scores are editorial rubric ratings synthesized from benchmark and feature evidence, not vendor benchmark results. Keep them separate from benchmark values."},"economics":{"effort_factors":{"none":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4},"cache_hit_presets":{"cold":0,"mixed":0.4,"warm":0.7,"hot":0.9},"models":[{"model_id":"fable-5","base_blended_price":38,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"],"note":"Always-on adaptive thinking; effort is a behavioral control, not a published token multiplier."},{"model_id":"gpt-5.6-sol","base_blended_price":22.5,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-terra","base_blended_price":9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-luna","base_blended_price":0.9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"opus-4.8","base_blended_price":19,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"]},{"model_id":"kimi-k3","base_blended_price":11.4,"cache_read_ratio":0.1,"efforts":["max"],"note":"Always-on reasoning at launch; the max-only effort control is behavioral and not a published token multiplier."}],"efficiency_filters":{"quality_floor_default":75,"provider_default":"all","effort_default":"medium","workload_default":"mixed","sort_default":"scq_compound_desc"}},"hardware_options":[{"id":"local-64gb","name":"64 GB unified-memory workstation","memory_gb":64,"type":"local","best_for":"3B-35B dense or small MoE models at Q4-Q8","cost_note":"Capex varies; compare measured memory bandwidth, not product year."},{"id":"local-128gb","name":"128 GB unified-memory workstation","memory_gb":128,"type":"local","best_for":"Up to roughly 100 GB quantized weights with headroom for KV cache","cost_note":"Portable and quiet; slower than datacenter GPUs for sustained batches."},{"id":"cloud-2xa6000","name":"2 x RTX A6000 48 GB","memory_gb":96,"type":"cloud","best_for":"70B-class dense and mid-size MoE Q4 deployments","cost_note":"Spot price varies by host; record price and interconnect at test time."},{"id":"cloud-h200","name":"1 x H200 141 GB","memory_gb":141,"type":"cloud","best_for":"High-throughput 70B inference and larger quantized MoE models","cost_note":"Use provider quote; hourly rates change frequently."},{"id":"cloud-b200","name":"1 x B200 192 GB","memory_gb":192,"type":"cloud","best_for":"Large-model throughput where software stack supports Blackwell","cost_note":"Availability and hourly rates vary; validate framework support."}],"hardware_model_fit":[{"model_id":"qwen3.6-35b-a3b","hardware_id":"local-64gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":55,"evidence":"planning_estimate"},{"model_id":"qwen3.6-35b-a3b","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"BF16","estimated_output_tps":140,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"local-64gb","fit_status":"unsupported","reason":"Quantized weights exceed usable memory."},{"model_id":"kimi-k2.6-1t","hardware_id":"local-128gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":16,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"cloud-h200","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":36,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-flash-open","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":45,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-h200","fit_status":"unsupported","reason":"Single-device memory is insufficient at the tracked quantization."},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-b200","fit_status":"supported","quant":"Q2_K","estimated_output_tps":14,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"local-64gb","fit_status":"supported","quant":"BF16","estimated_output_tps":38,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"cloud-h200","fit_status":"supported","quant":"BF16","estimated_output_tps":80,"evidence":"planning_estimate"}],"decision_support":{"recommended_stack":[{"route":"hardest_long_horizon","primary":"fable-5","effort":"high","fallback":"gpt-5.6-sol","why":"Use Fable where long-horizon reliability warrants cost and retention policy is acceptable."},{"route":"default_engineering","primary":"gpt-5.6-terra","effort":"medium","fallback":"gpt-5.6-sol","why":"Best balanced default in the current benchmark/cost set."},{"route":"high_volume_subagents","primary":"gpt-5.6-luna","effort":"low","fallback":"gemini-3.5-flash-lite","why":"Luna retains stronger coding-agent evidence; Flash-Lite is the throughput-first fallback for extraction and delegated subtasks."},{"route":"privacy_local","primary":"qwen3.6-35b-a3b","effort":null,"fallback":"deepseek-v4-flash-open","why":"Local-first route where external retention is unacceptable."},{"route":"vision_long_context","primary":"gemini-3.1-pro","effort":"high","fallback":"gpt-5.6-sol","why":"Prefer for multimodal context; validate preview stability before production."}],"routing_rules":["Route by task risk and evidence, not provider family.","Start bulk work on Luna or Terra and escalate only after an explicit verification failure.","Do not send ZDR-required data to Fable 5 because the model requires 30-day retention.","Record model, effort, cache state, region and harness version for every internal speed comparison.","Use self-hosted routes only when the selected hardware row is supported; never infer performance from an empty cell."]},"actions":[{"priority":"P0","owner":"Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","status":"open","order":1},{"priority":"P0","owner":"Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","status":"open","order":2},{"priority":"P0","owner":"Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","status":"open","order":3},{"priority":"P1","owner":"Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","status":"open","order":4},{"priority":"P1","owner":"Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","status":"open","order":5},{"priority":"P2","owner":"Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics.","status":"open","order":6}],"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the price-change history using OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. Added the paid Codex and ChatGPT Work credit reduction, unchanged subscription prices and quota budgets, and Sol API Fast mode (up to 2.5× Standard speed at 2× price)."},{"date":"2026-08-01","tag":"pricing","text":"Refreshed OpenAI's live GPT-5.6 API rate card: Terra is now $2/$12 per million input/output tokens and Luna is $0.20/$1.20, with cached input at $0.20 and $0.02 respectively. Recomputed the 30/70 workload blend, cost-efficiency scores, SCQ compounds and burn baselines. Superseded on August 7 with the official July 30 effective date and full change details."},{"date":"2026-08-01","tag":"benchmark","text":"Recorded DeepSeek V4-Flash-0731's July 31 public-beta evidence: Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in DeepSeek Harness minimal mode. Additional vendor-reported NL2Repo, Cybergym, Toolathlon and Automation Bench results remain narrative evidence rather than new shared columns."},{"date":"2026-07-22","tag":"method","text":"Marked the open-weight quality history as a Fable-anchored source series that is rebased in the browser against the selected reference; fixed task-speed remains explicitly Fable-indexed."},{"date":"2026-07-22","tag":"model","text":"Added Gemini 3.6 Flash and Gemini 3.5 Flash-Lite benchmark, pricing, context and throughput evidence from Google and Artificial Analysis; retained missing task-time fields as null."},{"date":"2026-07-21","tag":"model","text":"Added official Kimi K3 API pricing and derived the documented 30/70 workload blend and cost-efficiency score; no task-speed compound is shown because comparable speed evidence remains unavailable."},{"date":"2026-07-19","text":"Added a dated GPU-hosting price tracker with billing-basis normalization, historical Runpod price changes, current Contabo configurations, and quote-aware Infomaniak options. Replaced stale Vast.ai constants with live-marketplace status."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, context, benchmark records and reference eligibility."},{"date":"2026-07-18","tag":"method","text":"Introduced documented quality, speed, cost, SCQ and six-axis capability composites with explicit missing-data rules."},{"date":"2026-07-18","tag":"hardware","text":"Replaced blank hardware fit speeds with supported/unsupported states and evidence-labelled planning estimates."},{"date":"2026-07-18","tag":"routing","text":"Updated recommended routing to Fable for hardest retained-data work, Terra for default engineering and Luna for high-volume subagents."}],"sources":[{"id":"openai-gpt-5-6","title":"GPT-5.6 launch and evaluations","url":"https://openai.com/index/gpt-5-6/","accessed":"2026-07-18"},{"id":"moonshot-kimi-k3-pricing","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing","accessed":"2026-07-21"},{"id":"openai-models","title":"OpenAI model catalog","url":"https://developers.openai.com/api/docs/models","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-terra","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-luna","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-price-performance","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","accessed":"2026-08-07"},{"id":"openai-gpt-5-6-chat-access","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","accessed":"2026-08-07"},{"id":"deepseek-v4-flash-0731","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","accessed":"2026-08-01"},{"id":"deepseek-v4-pricing","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/","accessed":"2026-08-01"},{"id":"anthropic-models","title":"Claude models overview","url":"https://platform.claude.com/docs/en/about-claude/models/overview","accessed":"2026-07-18"},{"id":"anthropic-retention","title":"Claude API and data retention","url":"https://platform.claude.com/docs/en/manage-claude/api-and-data-retention","accessed":"2026-07-18"},{"id":"google-gemini-3-5","title":"Gemini 3.5 Flash","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash","accessed":"2026-07-18"},{"id":"google-gemini-3-6","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash","accessed":"2026-07-22"},{"id":"google-gemini-3-5-flash-lite","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite","accessed":"2026-07-22"},{"id":"google-flash-july-2026","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/","accessed":"2026-07-22"},{"id":"aa-gemini-3-6-flash","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash","accessed":"2026-07-22"},{"id":"aa-gemini-3-5-flash-lite","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","accessed":"2026-07-22"},{"id":"moonshot-kimi-k3","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3","accessed":"2026-07-18"},{"id":"moonshot-models","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models","accessed":"2026-07-18"},{"id":"openai-gpt-5-5","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/","accessed":"2026-07-18"}],"hosting_prices":{"as_of":"2026-07-19","hours_per_month":730,"methodology":"Preserve provider currency and billing basis. normalized_hourly is a comparison-only derivation for fixed monthly plans; monthly_equivalent multiplies hourly rates by 730. It excludes tax, storage, egress, public IPs, support, discounts, utilization, and model throughput. Append observations instead of replacing them.","trend_policy":"Track the same provider, GPU, service tier, region basis, and currency. A changed SKU starts a new series. Two or more observations produce a trend; one observation is a dated baseline.","offers":[{"id":"runpod-h100-sxm-secure","provider":"Runpod","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$2.99/hr","normalized_hourly":2.99,"monthly_equivalent":2182.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":3.99,"source_n":64},{"date":"2026-07-19","value":2.99,"source_n":63}]},{"id":"runpod-a100-sxm-secure","provider":"Runpod","configuration":"1 x A100 SXM","gpu":"A100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$1.49/hr","normalized_hourly":1.49,"monthly_equivalent":1087.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.94,"source_n":64},{"date":"2026-07-19","value":1.49,"source_n":63}]},{"id":"runpod-l40s-secure","provider":"Runpod","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$0.99/hr","normalized_hourly":0.99,"monthly_equivalent":722.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.19,"source_n":64},{"date":"2026-07-19","value":0.99,"source_n":63}]},{"id":"hyperstack-h100-sxm","provider":"Hyperstack","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$2.40/hr","normalized_hourly":2.4,"monthly_equivalent":1752,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":2.4,"source_n":68}]},{"id":"hyperstack-h200-sxm","provider":"Hyperstack","configuration":"1 x H200 SXM","gpu":"H200 SXM","gpu_count":1,"vram_gb":141,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$3.50/hr","normalized_hourly":3.5,"monthly_equivalent":2555,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":3.5,"source_n":68}]},{"id":"contabo-l40s-monthly","provider":"Contabo","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€751/mo","normalized_hourly":1.0288,"monthly_equivalent":751,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":1.0288,"source_n":65}]},{"id":"contabo-h100-monthly","provider":"Contabo","configuration":"1 x H100","gpu":"H100","gpu_count":1,"vram_gb":80,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€1,838/mo","normalized_hourly":2.5178,"monthly_equivalent":1838,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.5178,"source_n":65}]},{"id":"contabo-h200-nvl-monthly","provider":"Contabo","configuration":"1 x H200 NVL","gpu":"H200 NVL","gpu_count":1,"vram_gb":141,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€2,149/mo","normalized_hourly":2.9438,"monthly_equivalent":2149,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.9438,"source_n":65}]},{"id":"infomaniak-l40s","provider":"Infomaniak","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Switzerland","billing":"Usage-based Public Cloud","price_basis":"Live calculator / availability validation","currency":"CHF","current_price_label":"Calculator / validate","normalized_hourly":null,"monthly_equivalent":null,"price_status":"not_publicly_exposed","checked":"2026-07-19","source_n":66,"history":[]},{"id":"vast-a6000-market","provider":"Vast.ai","configuration":"1 x RTX A6000","gpu":"RTX A6000","gpu_count":1,"vram_gb":48,"region":"Marketplace","billing":"Per second; on-demand, reserved, or interruptible","price_basis":"Live host marketplace","currency":"USD","current_price_label":"Live marketplace","normalized_hourly":null,"monthly_equivalent":null,"price_status":"dynamic_marketplace","checked":"2026-07-19","source_n":69,"history":[]}]}},"model_roster":{"schema_version":"2.0-preview","researched_at":"2026-08-07","default_reference":"fable-5","inclusion_policy":{"default_limit_per_provider":3,"default_scope":"core","summary":"Show a flagship, a balanced or fast model, and one distinctive specialist per provider. Keep siblings and emerging models in expandable extended and watchlist scopes.","promotion_rule":"Promote a watchlist family when at least two signals hold: current official release, meaningful adoption or discussion, differentiated capability or efficiency, and reproducible availability."},"speed_methodology":{"dimensions":["time_to_first_token","output_tokens_per_second","end_to_end_task_time"],"rule":"Compare speed only at the exact model, configuration, provider, region, date, workload, and cache condition. Vendor claims are labelled and unknown values remain unknown.","bands":{"fast":"Designed or measured for low interaction latency","balanced":"General-purpose latency and quality trade-off","deliberate":"Higher reasoning depth or multi-agent execution","unknown":"No comparable evidence"}},"models":[{"id":"fable-5","name":"Claude Fable 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","reference":true,"context":"1M","price":"$10 / $50","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"deliberate","speed_note":"Anthropic comparative latency: slower. No Fable fast mode.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"30-day retention; no zero-data-retention option","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"opus-5","name":"Claude Opus 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Anthropic comparative latency: moderate. Research-preview fast mode claims up to 2.5x higher output throughput at $10 / $50 per MTok.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"Adaptive thinking is on by default; disabling it is supported only through high effort. Independent benchmark results vary materially by effort level and should not be treated as a single generic score.","source":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5"},{"id":"opus-4.8","name":"Claude Opus 4.8","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Optional API fast mode claims up to 2.5x output throughput at premium pricing.","availability":"GA","source":"https://platform.claude.com/docs/en/build-with-claude/fast-mode"},{"id":"sonnet-5","name":"Claude Sonnet 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$3 / $15","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"fast","speed_note":"Anthropic comparative latency: fast.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"haiku-4.5","name":"Claude Haiku 4.5","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"200K","price":"$1 / $5","license":"Proprietary","control":"fixed","levels":[],"default_level":"n/a","speed_band":"fast","speed_note":"Anthropic’s latest verified Haiku and fastest listed Claude; no Haiku 5 is in the current official catalog.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports max reasoning completes comparable AAII work in 61% less time than Fable 5. API Fast mode, introduced July 30, claims up to 2.5x Standard speed at 2x price with unchanged intelligence; no comparable output-tokens/s figure is published.","availability":"GA; ChatGPT, Codex, API","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$2 / $12; cached input $0.20","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 20% from $2.50/$15; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; subscription prices and quota budgets unchanged","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$0.20 / $1.20; cached input $0.02","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"fast","speed_note":"OpenAI positions Luna as the fastest tier and reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 80% from $1/$6; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; rolling out as Free and Go default with unlimited text chats and a Think option","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.5","name":"GPT-5.5","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M API / 400K Codex","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports GPT-5.4-class per-token latency; Codex fast mode is 1.5x throughput at 2.5x cost.","availability":"API, ChatGPT and Codex","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"extended","status":"stable","context":"1M","price":"$30 / $180","license":"Proprietary","control":"reasoning effort","levels":["medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Higher-accuracy tier for difficult work; no comparable public output-throughput figure.","availability":"API and eligible ChatGPT plans","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"harness_bundled","scope":"extended","status":"stable","context":"400K","price":"Included by plan / routed service","license":"Proprietary","control":"fixed","levels":[],"default_level":"provider default","speed_band":"fast","speed_note":"Low-latency ChatGPT route; do not equate service behavior with the reasoning model API.","availability":"ChatGPT","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"preview","context":"1M","price":"$2 / $12","license":"Proprietary","control":"thinking level","levels":["low","medium","high"],"default_level":"high","speed_band":"deliberate","speed_note":"Google warns high thinking may significantly delay the first answer token.","availability":"Preview","source":"https://ai.google.dev/gemini-api/docs/thinking"},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$1.50 / $7.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Artificial Analysis measured about 304 output tok/s at high thinking; Google reports fewer reasoning turns and tool calls than 3.5 Flash.","availability":"GA via Gemini API, AI Studio, Gemini app and Antigravity","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$1.50 / $9","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Still offered; migrate new general Flash workloads to 3.6 Flash for lower output price and stronger agentic performance.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash"},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$0.30 / $2.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Artificial Analysis measured about 490 output tok/s; optimized for high-volume subagents, document parsing and extraction.","availability":"GA via Gemini API, AI Studio and Gemini app rollout","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$0.25 / $1.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Still offered; Gemini 3.5 Flash-Lite is the current high-throughput migration target.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite"},{"id":"grok-4.5","name":"Grok 4.5","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"500K","price":"$2 / $6","license":"Proprietary","control":"reasoning effort","levels":["low","medium","high"],"default_level":"high","speed_band":"fast","speed_note":"xAI reports 80 output tokens/s; vendor measurement conditions are not fully comparable here.","availability":"API and integrations; verify regional availability","source":"https://x.ai/news/grok-4-5"},{"id":"grok-4.20-multi-agent","name":"Grok 4.20 Multi-Agent","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"extended","status":"specialist","context":"1M","price":"$1.25 / $2.50","license":"Proprietary","control":"agent count","levels":["low","medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Levels control a 4- or 16-agent process, not ordinary reasoning depth.","availability":"API","source":"https://docs.x.ai/developers/model-capabilities/text/multi-agent"},{"id":"composer-2.5","name":"Composer 2.5","family":"Composer","provider":"Cursor","region":"US","channel":"harness_bundled","scope":"core","status":"stable","context":"Managed by Cursor","price":"$0.50 / $2.50 standard","license":"Proprietary","control":"speed variant","levels":["standard","fast"],"default_level":"fast","speed_band":"fast","speed_note":"Fast is the default; Cursor documents no public low/medium/high effort selector.","availability":"Cursor only","source":"https://cursor.com/blog/composer-2-5"},{"id":"kimi-k3","name":"Kimi K3","family":"Kimi K3","provider":"Moonshot AI","region":"China","channel":"commercial_api_open_weights_announced","scope":"core","status":"stable","context":"1M","price":"$3 / $15; cached input $0.30","license":"Terms pending weight release","control":"reasoning effort","levels":["max"],"default_level":"max","speed_band":"deliberate","speed_note":"No comparable K3 output-rate figure at launch; max thinking is always enabled.","availability":"API, Kimi, Kimi Work and Kimi Code; weights promised by July 27","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"id":"kimi-k2.7-code","name":"Kimi K2.7 Code","family":"Kimi K2.7","provider":"Moonshot AI","region":"China","channel":"commercial_api","scope":"core","status":"specialist","context":"256K","price":"Current API pricing","license":"Proprietary API","control":"service variant","levels":["standard","high-speed"],"default_level":"standard","speed_band":"fast","speed_note":"High-speed service is documented at about 180 tok/s and up to 260 tok/s for short contexts.","availability":"Kimi API and Kimi Code","source":"https://platform.kimi.ai/docs/models"},{"id":"deepseek-v4","name":"DeepSeek V4","family":"DeepSeek V4","provider":"DeepSeek","region":"China","channel":"open_weight","scope":"core","status":"public_beta","context":"1M","price":"Flash $0.14 / $0.28; cached input $0.0028","license":"MIT","control":"mode","levels":["non-think","think","max"],"default_level":"think","speed_band":"balanced","speed_note":"V4-Flash-0731 reports 82.7 on Terminal-Bench 2.1 at max effort in DeepSeek Harness minimal mode; comparable latency remains unknown.","availability":"V4-Flash-0731 API public beta and weights; future 2x peak-hours pricing announced without an effective date","variants":["Pro","Flash"],"source":"https://api-docs.deepseek.com/updates/"},{"id":"qwen-3.6","name":"Qwen3.6","family":"Qwen3.6","provider":"Alibaba Qwen","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"mode and token budget","levels":["non-thinking","thinking","budget"],"default_level":"thinking","speed_band":"balanced","speed_note":"Budget is provider-native; do not translate it into invented effort labels.","availability":"API and weights","variants":["27B","35B-A3B"],"source":"https://github.com/QwenLM/Qwen3.6"},{"id":"glm-5.2","name":"GLM-5.2","family":"GLM-5","provider":"Z.ai","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"MIT","control":"effort","levels":["adaptive"],"default_level":"adaptive","speed_band":"unknown","speed_note":"Architecture efficiency claims are not a comparable API latency measurement.","availability":"API and weights","source":"https://z.ai/blog/glm-5.2"},{"id":"kimi-k2.5","name":"Kimi K2.5","family":"Kimi K2","provider":"Moonshot AI","region":"China","channel":"open_weight","scope":"extended","status":"deprecated","context":"256K","price":"API and self-hosted","license":"Modified MIT","control":"mode","levels":["instant","thinking"],"default_level":"thinking","speed_band":"balanced","speed_note":"No comparable official throughput measurement; Instant is the lower-latency mode.","availability":"Existing users only; platform sunset August 31, 2026","source":"https://platform.kimi.ai/docs/models"},{"id":"minimax-m3","name":"MiniMax M3","family":"MiniMax M3","provider":"MiniMax","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"Community; non-commercial by default","control":"thinking mode","levels":["disabled","adaptive","enabled"],"default_level":"adaptive","speed_band":"balanced","speed_note":"Vendor claims decode improvements versus M2, not cross-provider latency.","availability":"API and weights","caveat":"License is not permissive open source","source":"https://www.minimax.io/blog/minimax-m3"},{"id":"mistral-small-4","name":"Mistral Small 4","family":"Mistral 3/4","provider":"Mistral AI","region":"EU","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"reasoning effort","levels":["none","high"],"default_level":"none","speed_band":"fast","speed_note":"Vendor claims 40% lower completion time and 3x requests/s versus Small 3.","availability":"API and weights","source":"https://mistral.ai/it/news/mistral-small-4/"},{"id":"llama-4","name":"Llama 4","family":"Llama 4","provider":"Meta","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"10M advertised for Scout","price":"Self-hosted","license":"Llama 4 Community License","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Performance depends on deployment; no comparable official API latency.","availability":"Weights","variants":["Scout","Maverick"],"caveat":"Not OSI-open; EU multimodal and large-platform restrictions apply","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"id":"gemma-3","name":"Gemma 3","family":"Gemma 3","provider":"Google","region":"US","channel":"open_weight","scope":"extended","status":"stable","context":"128K","price":"Self-hosted","license":"Google Gemma Terms","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Deployment dependent; optimized for edge and local use.","availability":"Weights","variants":["1B","4B","12B","27B"],"source":"https://ai.google.dev/gemma/docs/core/model_card_3"},{"id":"gpt-oss","name":"gpt-oss","family":"gpt-oss","provider":"OpenAI","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"128K","price":"Self-hosted","license":"Apache-2.0 plus usage policy","control":"reasoning effort","levels":["low","medium","high"],"default_level":"medium","speed_band":"unknown","speed_note":"Depends on hardware and serving stack.","availability":"Weights","variants":["120b","20b"],"source":"https://openai.com/index/introducing-gpt-oss/"},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","family":"Apertus v1.1","provider":"Swiss AI Initiative","region":"Switzerland","channel":"open_weight","scope":"core","status":"stable","context":"4K","price":"Self-hosted","license":"Apache-2.0","control":"sampling parameters","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"No comparable throughput measurement is published; performance depends on hardware, runtime and quantization.","availability":"Weights; Transformers, vLLM, SGLang and quantized checkpoints","variants":["0.5B","1.5B","4B"],"source":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"id":"inkling","name":"Inkling","family":"Inkling","provider":"Thinking Machines Lab","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"No metered API; open weights + Tinker fine-tuning platform","license":"Apache 2.0","control":"n/a","levels":[],"default_level":null,"speed_band":"unknown","speed_note":"No published throughput figure; not yet independently benchmarked for this roster.","availability":"Open weights (Hugging Face), Tinker fine-tuning platform, third-party inference providers","source":"https://thinkingmachines.ai/model-card/inkling/"}],"watchlist":[{"name":"Step 3.5 Flash","provider":"StepFun","region":"China","license":"Apache-2.0","signal":"Official 100-300 output tok/s claim","source":"https://github.com/stepfun-ai/Step-3.5-Flash"},{"name":"MiMo-V2-Flash","provider":"Xiaomi","region":"China","license":"Apache-2.0","signal":"Thinking toggle and generation-speed claim","source":"https://github.com/XiaomiMiMo/MiMo-V2-Flash"},{"name":"Seed-OSS-36B","provider":"ByteDance","region":"China","license":"Apache-2.0","signal":"512K context and controllable thinking budget","source":"https://github.com/ByteDance-Seed/seed-oss"},{"name":"Hy3 Preview","provider":"Tencent","region":"China","license":"Tencent Hy Community License","signal":"Preview multimodal MoE family","source":"https://github.com/Tencent-Hunyuan/Hy3-preview"}]}}&lt;/script&gt;
&lt;section class="tab-panel" id="tab-strategy"&gt;
 &lt;div data-section-sources="strategy"&gt;&lt;/div&gt;
 &lt;div id="strategy-content"&gt;&lt;/div&gt;
 &lt;/section&gt;
&lt;/div&gt;</description></item><item><title>Research Data Policy</title><link>https://projectious-work.github.io/ai-market-research/docs/research-data-policy/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/docs/research-data-policy/</guid><description>&lt;p&gt;This policy applies to market facts, benchmark observations, pricing,
availability, source links, and historical snapshots used by the Signal Room.&lt;/p&gt;
&lt;h2 id="source-rights-and-collection"&gt;Source rights and collection&lt;/h2&gt;
&lt;p&gt;Use sources that are publicly accessible and lawful to consult. Prefer primary
sources such as official model cards, documentation, pricing pages, release
notes, repositories, and benchmark publications. Do not bypass access controls,
paywalls, authentication, rate limits, or technical restrictions.&lt;/p&gt;
&lt;p&gt;Store only the facts and short summaries needed for market analysis. Do not
mirror articles, proprietary datasets, benchmark submissions, or other
copyrighted source material. A public URL does not imply permission to copy the
underlying work.&lt;/p&gt;</description></item><item><title>05 Evidence</title><link>https://projectious-work.github.io/ai-market-research/report/05-evidence/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/report/05-evidence/</guid><description>&lt;div class="sr-report-scope"&gt;
&lt;script id="market-data" type="application/json"&gt;{"meta":{"generated_at":"2026-08-07T00:00:00Z","reference_default":"fable-5","report_metrics_file":"data/report-metrics.json"},"executive_summary":{"models":["**Claude Opus 5 is now generally available.** Anthropic's new `claude-opus-5` targets complex agentic coding and enterprise work with a 1M-token context window, 128K maximum output, adaptive thinking by default, and unchanged $5/$25 per MTok base pricing. Independent benchmark values are not yet recorded here.","**Google reset the Flash price-performance curve on July 21.** Gemini 3.6 Flash is GA at $1.50/$7.50 per million tokens with stronger coding and agentic results than 3.5 Flash, while Gemini 3.5 Flash-Lite reaches roughly 490 output tok/s at $0.30/$2.50 for high-volume subagents and extraction.","**Kimi K3 resets the open-weight frontier.** Moonshot reports a 2.8T sparse MoE, native vision, 1M context, and 99% of Fable 5 across 14 overlapping launch-suite evaluations; weights are promised by July 27.","**GPT-5.6 Luna and Terra became materially cheaper on July 30.** Luna fell 80% to $0.20/$1.20 per MTok and Terra 20% to $2/$12; their paid Codex and ChatGPT Work usage also consumes fewer credits, while subscription prices and quota budgets did not change.","**GPT-5.6 now spans Sol, Terra and Luna**, giving leaders a deliberate capability, balanced, and high-throughput ladder under one family. Sol API Fast mode replaces Priority Processing: OpenAI claims up to 2.5× Standard speed at 2× price, with unchanged intelligence.","**Luna is expanding beyond API routing.** OpenAI says it becomes the default for Free and Go users, with unlimited text chats and a higher-reasoning Think option rolling out subject to abuse guardrails. This is a ChatGPT product update, not an API capability change.","**Keep benchmark tables separate by harness.** OpenAI's GPT-5.6 launch results and Scale's public SWE-bench Pro leaderboard use different model versions and evaluation setups; use each as evidence, not as one directly rankable series.","**GPT-5.5 remains active in Codex and the API.** It is retained as a compatibility and portfolio option rather than being hidden by the newer 5.6 family.","**Claude Haiku 4.5 remains Anthropic’s latest verified Haiku.** No official Haiku 5 listing was found in Anthropic’s current model catalog."],"harnesses":["**Separate model capability from harness capability.** Tool execution, context management, isolation, and observability can dominate real workflow outcomes.","**Subscription access is not a production routing contract.** Validate API, credit-pool, and third-party harness policies before standardizing an operating model.","**Maintain at least one portable fallback path** across providers for high-value workflows and operational incidents.","**Gemini Managed Agents now add background execution, remote MCP, custom functions and credential refresh.** That makes Google a more credible managed-agent control plane, but it does not substitute for workload-specific evaluation."],"self_hosting":["**Open weights are now a strategic option, not only a cost play.** Kimi K3, DeepSeek, Qwen, Nemotron, Mistral, and Llama cover different sovereignty and specialization needs.","**Do not compare self-hosting at zero token cost.** Include accelerator rental or depreciation, power, utilization, serving staff, and measured throughput.","**Pilot against a defined workload and hardware envelope** before treating advertised context or parameter scale as deployable capacity."],"strategy":["**Run a portfolio, not a winner-takes-all model standard.** Reserve frontier reasoning for high-value decisions and route routine work to measured fast or efficient tiers.","**Instrument quality, latency, retries, and total workflow cost together.** Token price alone is not an operating metric.","**Review the portfolio quarterly and after major releases**, with explicit retirement, security, and fallback criteria."]},"headline_stats":[{"id":"frontier_count","label":"Models tracked","value":35,"unit":"","delta":"Current, fast and retained compatibility models","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_weight","label":"Open-weight","color":"#16866f"}],"history":[{"date":"2026-05-18","proprietary":20,"open_weight":3},{"date":"2026-07-22","proprietary":26,"open_weight":8},{"date":"2026-07-25","proprietary":27,"open_weight":8}],"note":"Release-tag roster snapshots; announced open-weight releases are counted with open-weight models."}},{"id":"best_open_pct","label":"Best open-weight vs selected reference","value":99,"unit":"%","delta":"Kimi K3; geometric mean across 14 overlapping vendor evals","delta_dir":"up","signal_id":"open_weight_quality"},{"id":"cheapest_frontier_api","label":"Lowest frontier blended API price","value":0.9,"unit":"$/Mtok","delta":"GPT-5.6 Luna; 30% input / 70% output blend","delta_dir":"down","signal_id":"blended_frontier_price"},{"id":"fastest_task_rate","label":"Fastest task-rate index","value":"3.2×","unit":"","delta":"Fable 5 = 1.0×; planning index, not tokens/second","delta_dir":"up","signal_id":"task_speed"},{"id":"max_context","label":"Largest usable context","value":"10M","unit":"tok","delta":"Llama 4 Scout; deployment constraints still apply","delta_dir":"up","trend":[{"date":"2026-01-01","value":1},{"date":"2026-04-01","value":1},{"date":"2026-07-01","value":10}]},{"id":"harness_count","label":"Harnesses tracked","value":16,"unit":"","delta":"Commercial and open agent environments","delta_dir":"up","stacked_trend":{"series":[{"key":"proprietary","label":"Proprietary","color":"#e05232"},{"key":"open_source","label":"Open source","color":"#3d78c5"}],"history":[{"date":"2026-05-18","proprietary":3,"open_source":11},{"date":"2026-07-22","proprietary":6,"open_source":10}],"note":"Release-tag roster snapshots classified from each harness license."}},{"id":"documented_output_speed","label":"Fastest documented API output","value":490,"unit":"tok/s","delta":"Gemini 3.5 Flash-Lite; Artificial Analysis first-party API measurement","delta_dir":"up","signal_id":"output_throughput"},{"id":"quality_coverage","label":"Models with quality evidence","value":"31/35","unit":"","delta":"Unknown remains unknown; no zero-value substitution","delta_dir":"neutral","trend":[{"date":"2026-01-01","value":18},{"date":"2026-04-01","value":24},{"date":"2026-07-01","value":31}]}],"trends":{"best_open_vs_opus":{"label":"Best open-weight % vs Opus 4.7","history":[{"date":"2025-11-01","value":68},{"date":"2025-12-01","value":72},{"date":"2026-01-01","value":78},{"date":"2026-02-01","value":80},{"date":"2026-03-01","value":83},{"date":"2026-04-01","value":86},{"date":"2026-05-01","value":88},{"date":"2026-06-01","value":90}]},"median_frontier_output_price":{"label":"Median frontier $/Mtok output","history":[{"date":"2025-11-01","value":18},{"date":"2025-12-01","value":17},{"date":"2026-01-01","value":15.5},{"date":"2026-02-01","value":14.5},{"date":"2026-03-01","value":13},{"date":"2026-04-01","value":12.5},{"date":"2026-05-01","value":12},{"date":"2026-06-01","value":11}]},"models_released_per_month":{"label":"Notable model releases per month","history":[{"date":"2025-11-01","value":3},{"date":"2025-12-01","value":4},{"date":"2026-01-01","value":5},{"date":"2026-02-01","value":6},{"date":"2026-03-01","value":4},{"date":"2026-04-01","value":7},{"date":"2026-05-01","value":5},{"date":"2026-06-01","value":8}]}},"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the GPT-5.6 pricing history from OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. The report now records lower paid Codex and ChatGPT Work credit consumption, unchanged subscription prices and quota budgets, Sol API Fast mode (up to 2.5× Standard speed at 2× price), and the August 6 ChatGPT Luna access expansion."},{"date":"2026-08-01","tag":"pricing","text":"Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to $0.20/$1.20 from OpenAI's live model pages. Recomputed the tracked 30/70 workload blends and effort-burn ratios. Superseded on August 7 with OpenAI's published July 30 effective date and full price-change details."},{"date":"2026-08-01","tag":"model","text":"Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output, thinking/non-thinking modes, $0.14/$0.28 per-million pricing, $0.0028 cached input, and vendor-reported agent results including Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2."},{"date":"2026-07-28","tag":"model","text":"Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B active) open-weight (Apache 2.0) multimodal MoE with 1M-token context, released 2026-07-15. No first-party per-token API pricing is published (monetized via the Tinker fine-tuning platform); benchmark scores are published on the model card but not yet normalized into this roster's comparable set, so quality/speed/cost fields are left unknown rather than estimated."},{"date":"2026-07-25","tag":"model","text":"Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive thinking by default, $5/$25 per MTok base pricing, and official cloud-platform availability. Added independent Artificial Analysis evidence (61 Intelligence Index at max effort; 52.3 output tok/s) with effort-specific caveats. Gemini 3.5 Flash Cyber remains limited to CodeMender government and trusted-partner pilots; GPT-Live and Muse Spark 1.1 remain non-API products, so none were added to the API roster. Claude Opus 4.7 Fast Mode was removed July 24."},{"date":"2026-07-24","tag":"benchmark","text":"Refreshed current-source evidence: added OpenAI's GPT-5.6 launch table, Google's Managed Agents update, and Scale's public SWE-bench Pro leaderboard. Clarified that vendor launch tables and the public leaderboard are not directly comparable because their model versions and harnesses differ."},{"date":"2026-07-22","tag":"fix","text":"Made the selected reference propagate through open-weight quality headlines, market-signal history, comparison headings, model analytics, self-hosting quality and capability market position. Replaced the model and harness inventory mini-lines with stacked proprietary/open category areas based on release-tag roster snapshots."},{"date":"2026-07-22","tag":"model","text":"Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M context, 64K output, current API pricing, Artificial Analysis intelligence and throughput measurements, and Google’s published coding and agentic benchmarks."},{"date":"2026-07-22","tag":"data","text":"Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to governments and trusted CodeMender partners; Gemini Omni Flash and Nano Banana 2 Lite remain specialized media models rather than general-purpose roster entries."},{"date":"2026-07-21","tag":"model","text":"Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint: fully open Apache 2.0 weights and data, 4K context, 1.7T-token distillation, 1,811 languages, and official BF16, FP8, NVFP4A16, INT3, INT4 and INT6 variants."},{"date":"2026-07-21","tag":"harness","text":"Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI 0.144.6, Gemini CLI 0.51.0, OpenCode 1.18.4, Cline 4.0.10, and Goose 1.43.0; updated repository star snapshots and notable release capabilities."},{"date":"2026-07-21","tag":"model","text":"Added Moonshot's official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached input, and $15/M output; the 30/70 workload blend is $11.40/M before reasoning-effort effects."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark registers and selectable reference configurations. Added documented quality, speed, cost and capability composites in data/report-metrics.json."},{"date":"2026-07-18","tag":"routing","text":"Updated the action queue and recommended routing: Fable for hardest retained-data workloads, Terra for default engineering, Luna for high-volume subagents, with explicit escalation rules."},{"date":"2026-06-06","tag":"fix","text":"Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K context, 1000+ tok/s on Cerebras) and Hermes Agent v0.16.0 (Nous Research, MIT, self-hosted multi-platform agent) — both were incorrectly removed in v1.5.0 sweep."},{"date":"2026-06-06","tag":"policy","text":"Dashboard market sweep v1.5.0: real-world re-grounding. Replaced fictional Mythos/GPT-5.5-Cyber rows with verified models; added Nvidia Nemotron coalition, Kimi K2.6, GLM-5, Cohere Command A+, SubQ 1M-Preview."},{"date":"2026-06-04","tag":"model","text":"Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context, NVIDIA Open Model License) at Computex — first frontier-scale open model from Nvidia."},{"date":"2026-06-04","tag":"model","text":"Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral, Perplexity, Reflection AI, Sarvam, Thinking Machines Lab as inaugural members."},{"date":"2026-06-01","tag":"model","text":"Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model."}],"actions":["P0 · Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","P0 · Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","P0 · Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","P1 · Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","P1 · Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","P2 · Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics."],"models":[{"id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","tier":"frontier","released":"2026-06-09","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":80,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii":59.9,"coding_agent_index":77.2,"deep_swe":69.7,"terminal_bench":83.1,"agents_last_exam":40.5,"api_in":10,"api_out":50,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":52.3,"subscription":"Claude API and supported cloud platforms","notes":"Default reference for v2. Anthropic describes Fable 5 as its most capable widely released model. Adaptive thinking is always on. Comparative latency is slower. Retention is 30 days and zero-data-retention is not available.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"status":"stable","speed_class":"deliberate","speed_evidence":"vendor-qualitative","capability_levels":{"coding":49,"reasoning":50,"knowledge":48,"comms":48,"multimodal":45,"agentic":50}},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":64.6,"swe_verified":null,"aaii":58.9,"coding_agent_index":80,"deep_swe":72.7,"terminal_bench":88.8,"agents_last_exam":52.7,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Flagship GPT-5.6 tier. API Fast mode replaced Priority Processing on July 30: OpenAI claims up to 2.5× Standard speed at 2× Standard price with no intelligence change. Speed is otherwise stored as an end-to-end task index in report-metrics.json, not tok/s.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":49,"knowledge":48,"comms":48,"multimodal":48,"agentic":49}},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","tier":"frontier","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":63.4,"swe_verified":null,"aaii":55,"coding_agent_index":77.4,"deep_swe":69.6,"terminal_bench":87.4,"agents_last_exam":50.4,"api_in":2,"api_out":12,"api_cache_hit":0.2,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Balanced GPT-5.6 tier and recommended default engineering route. On July 30, API list pricing fell 20% from $2.50/$15 to $2/$12 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":48,"reasoning":46,"knowledge":46,"comms":46,"multimodal":45,"agentic":48}},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","tier":"fast","released":"2026-07-09","license":"proprietary","jurisdiction":"US","context":1050000,"context_label":"1.05M (272K pricing threshold)","swe_pro":62.7,"swe_verified":null,"aaii":51.2,"coding_agent_index":74.6,"deep_swe":67.2,"terminal_bench":84.7,"agents_last_exam":50.3,"api_in":0.2,"api_out":1.2,"api_cache_hit":0.02,"batch_discount":50,"tok_per_sec":null,"subscription":"ChatGPT, Codex and API","notes":"Fastest, most affordable GPT-5.6 tier; recommended for high-volume subagents with verification and escalation. On July 30, API list pricing fell 80% from $1/$6 to $0.20/$1.20 per MTok; paid Codex and ChatGPT Work usage also consumes fewer credits. Subscription prices and quota budgets did not change. ChatGPT is also rolling Luna out as the Free and Go default, with unlimited text chats and a Think option; that product change does not alter API routing.","reasoning_capable":true,"effort_default":"medium","effort_levels":["none","low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":46,"reasoning":44,"knowledge":43,"comms":45,"multimodal":43,"agentic":46}},{"id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","tier":"frontier","released":"2026-07-16","license":"open weights announced; terms pending weight release","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":null,"tok_per_sec":null,"subscription":"Kimi API, Kimi Code, Kimi Work","reasoning_capable":true,"effort_default":"max","effort_levels":["max"],"notes":"2.8T sparse MoE with 16/896 experts active, native vision and always-on thinking. API pricing is $3/M uncached input, $0.30/M cached input and $15/M output. Full weights promised by July 27, 2026.","quality_vs_fable":99.2,"quality_evidence":"Geometric mean across 14 overlapping values in Moonshot’s launch comparison; vendor-reported, max/xhigh settings.","capability_levels":{"coding":47,"reasoning":46,"knowledge":47,"comms":46,"multimodal":47,"agentic":48}},{"id":"opus-5","name":"Claude Opus 5","provider":"Anthropic","tier":"frontier","released":"2026-07-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":null,"subscription":"Claude API, Bedrock, Google Cloud, Microsoft Foundry","notes":"Current Anthropic Opus generation. 1M context and 128K maximum output. Adaptive thinking is enabled by default; effort defaults to high. Artificial Analysis reports a 61 Intelligence Index and 52.3 output tok/s at max effort; these figures are effort-specific. Research-preview Fast Mode is Claude API-only at $10/$50 per MTok and claims up to 2.5x higher output throughput.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"speed_class":"balanced","speed_evidence":"vendor-qualitative","cite":[86,87]},{"id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","tier":"frontier","released":"2026-05-28","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":69.2,"swe_verified":88.6,"livecodebench":82,"aime":90,"tau2":86,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Max 20× $200/mo · Max 5× $100/mo","notes":"Previous Opus generation, retained as an active compatibility option. SWE-V 88.6%, SWE-Pro 69.2%, AAII 61.4. Fast Mode reduced 3× to $10/$50 (was $30/$150 on 4.7). 1M context standard.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":50,"comms":50,"multimodal":34,"agentic":50}},{"id":"opus-4.7","name":"Claude Opus 4.7","provider":"Anthropic","tier":"frontier","released":"2026-04-16","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":64.3,"swe_verified":87.6,"livecodebench":79,"aime":88,"tau2":84,"api_in":5,"api_out":25,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":50,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Now legacy as of Opus 4.8 release May 28. Pricing unchanged. SWE-Verified 87.6%, SWE-Pro 64.3%. Fast Mode was removed July 24, 2026; standard-speed API access remains active.","reasoning_capable":true,"effort_default":"xhigh","effort_levels":["low","medium","high","xhigh","max"],"cache_discount":0.1,"capability_levels":{"coding":50,"reasoning":50,"knowledge":49,"comms":50,"multimodal":32,"agentic":50}},{"id":"sonnet-4.6","name":"Claude Sonnet 4.6","provider":"Anthropic","tier":"frontier","released":"2026-02-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":44,"swe_verified":77,"livecodebench":76,"aime":85,"tau2":81,"api_in":3,"api_out":15,"api_cache_hit":0.3,"batch_discount":50,"tok_per_sec":80,"subscription":"Max 5× $100/mo · Pro $20/mo","notes":"Best code style/intent understanding. With cache+batch: $0.30/$7.50 effective.","reasoning_capable":true,"effort_default":"high","effort_levels":["low","medium","high","max"],"cache_discount":0.1,"capability_levels":{"coding":41,"reasoning":40,"knowledge":39,"comms":42,"multimodal":31,"agentic":41}},{"id":"haiku-4.5","name":"Claude Haiku 4.5","provider":"Anthropic","tier":"fast","released":"2025-10-15","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":28,"swe_verified":62,"livecodebench":58,"aime":70,"tau2":65,"api_in":1,"api_out":5,"api_cache_hit":0.1,"batch_discount":50,"tok_per_sec":110,"subscription":"Available in all tiers","notes":"5× cheaper than Sonnet. Triage/classification champion.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.1,"capability_levels":{"coding":28,"reasoning":18,"knowledge":29,"comms":30,"multimodal":18,"agentic":28}},{"id":"gpt-5.5","name":"GPT-5.5","provider":"OpenAI","tier":"frontier","released":"2026-04-23","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M (272K threshold)","swe_pro":58.6,"swe_verified":88.7,"livecodebench":84,"aime":92,"tau2":82,"api_in":5,"api_out":30,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":55,"subscription":"Pro $200 · Pro Lite $100 · Plus $20","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"OpenAI flagship; default in ChatGPT (Instant variant since May 5, 2026). 1M context with 2× input/1.5× output surcharge above 272K. Reasoning tokens billed as output.","cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":41,"knowledge":49,"comms":40,"multimodal":41,"agentic":38}},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","provider":"OpenAI","tier":"frontier","released":"2026-04-24","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":60,"swe_verified":90,"livecodebench":86,"aime":94,"tau2":84,"api_in":30,"api_out":180,"api_cache_hit":3,"batch_discount":50,"tok_per_sec":35,"subscription":"Pro $200 only","reasoning_capable":true,"effort_default":"high","effort_levels":["medium","high","xhigh"],"notes":"Highest-stakes reasoning tier. $30/$180. Available in ChatGPT Pro $200 and as API. 6× cost of base 5.5.","cache_discount":0.25,"capability_levels":{"coding":47,"reasoning":48,"knowledge":50,"comms":42,"multimodal":42,"agentic":42}},{"id":"gpt-5.4","name":"GPT-5.4","provider":"OpenAI","tier":"frontier","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":57.7,"swe_verified":81,"livecodebench":81,"aime":90,"tau2":80,"api_in":2.5,"api_out":15,"api_cache_hit":0.25,"batch_discount":50,"tok_per_sec":60,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Best quality-per-credit. Held SWE-Pro lead Feb-April. Default for everyday coding.","cache_discount":0.1,"capability_levels":{"coding":40,"reasoning":40,"knowledge":40,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-mini","name":"GPT-5.4 Mini","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":38,"swe_verified":73,"livecodebench":71,"aime":78,"tau2":70,"api_in":0.4,"api_out":1.6,"api_cache_hit":0.04,"batch_discount":50,"tok_per_sec":130,"subscription":"All paid tiers","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"94% of GPT-5.4 coding at 6× less. Best subagent. ~1/20 burn vs GPT-5.5. Collapses at 64K+ context.","cache_discount":0.1,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":30,"agentic":30}},{"id":"gpt-5.4-nano","name":"GPT-5.4 Nano","provider":"OpenAI","tier":"fast","released":"2026-03-15","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":25,"swe_verified":60,"livecodebench":58,"aime":65,"tau2":55,"api_in":0.1,"api_out":0.4,"api_cache_hit":0.01,"batch_discount":50,"tok_per_sec":200,"subscription":"API only","reasoning_capable":true,"effort_default":"low","effort_levels":["minimal","low","medium"],"notes":"Smallest reasoning model. API-only. For embeddable/edge inference at near-zero cost.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":20,"comms":20,"multimodal":20,"agentic":20}},{"id":"gpt-5.3-codex","name":"GPT-5.3-Codex","provider":"OpenAI","tier":"frontier","released":"2026-01-20","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":56.8,"swe_verified":85,"livecodebench":82,"aime":87,"tau2":78,"api_in":1.5,"api_out":10,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":70,"subscription":"All paid tiers · Code Review uses this","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high","xhigh"],"notes":"Coding-specialised. ~⅓ burn vs GPT-5.5 for ~2pts less SWE-Pro. Often best $/quality for execution turns.","cache_discount":0.1,"capability_levels":{"coding":49,"reasoning":28,"knowledge":28,"comms":27,"multimodal":8,"agentic":38}},{"id":"gpt-5.3-codex-spark","name":"GPT-5.3-Codex-Spark","provider":"OpenAI","tier":"fast","released":"2026-02-12","license":"proprietary","jurisdiction":"US","context":128000,"context_label":"128K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":1000,"subscription":"ChatGPT Pro — research preview only","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"Smaller, latency-first sibling of GPT-5.3-Codex. Released Feb 12, 2026. 1000+ tok/s on Cerebras hardware. ChatGPT Pro research preview only — not in API at launch; separate preview rate-limit pool (no standard credit burn). Text-only. Target use: real-time micro-edits, live pair-programming in Codex app/CLI/VS Code.","cache_discount":null,"capability_levels":{"coding":38,"reasoning":22,"knowledge":22,"comms":25,"multimodal":0,"agentic":28}},{"id":"gpt-5.2","name":"GPT-5.2","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":52,"swe_verified":78,"livecodebench":76,"aime":84,"tau2":75,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Available but not recommended","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Codex picker keeps it for 'long-running agents' (specifically tuned for autonomy). Otherwise eclipsed by 5.3-Codex.","cache_discount":0.1},{"id":"gpt-5.2-codex","name":"GPT-5.2-Codex","provider":"OpenAI","tier":"legacy","released":"2025-11-10","license":"proprietary","jurisdiction":"US","context":200000,"context_label":"200K","swe_pro":50,"swe_verified":76,"livecodebench":75,"aime":82,"tau2":73,"api_in":1.25,"api_out":8,"api_cache_hit":0.125,"batch_discount":50,"tok_per_sec":65,"subscription":"Legacy","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"notes":"Predecessor to 5.3-Codex. Legacy.","cache_discount":0.1},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","provider":"OpenAI","tier":"fast","released":"2026-05-05","license":"proprietary","jurisdiction":"US","context":400000,"context_label":"400K","swe_pro":35,"swe_verified":76,"livecodebench":72,"aime":81,"tau2":70,"api_in":1.5,"api_out":6,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":200,"subscription":"Default ChatGPT model · API chat-latest","reasoning_capable":false,"effort_default":null,"effort_levels":[],"notes":"ChatGPT default since May 5, 2026. 52.5% fewer hallucinations vs 5.3, 30% shorter responses, first Instant-class High-capability rating on cybersec/bio-chem.","cache_discount":0.1,"capability_levels":{"coding":20,"reasoning":20,"knowledge":40,"comms":30,"multimodal":30,"agentic":10}},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","provider":"Google","tier":"frontier","released":"2026-02-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":54.2,"swe_verified":78,"livecodebench":76,"aime":85,"tau2":78,"api_in":2,"api_out":12,"api_cache_hit":0.5,"batch_discount":50,"tok_per_sec":119,"subscription":"Gemini Advanced $20 · Ultra $100","notes":"Released Feb 19, 2026 (preview). Pricing doubles above 200K input tokens. 50% batch discount.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":39,"reasoning":41,"knowledge":49,"comms":41,"multimodal":50,"agentic":29}},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":58.7,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78,"agents_last_exam":null,"quality_vs_fable":81.6,"api_in":1.5,"api_out":7.5,"api_cache_hit":0.15,"batch_discount":50,"tok_per_sec":304,"subscription":"Gemini API · AI Studio · Gemini app · Antigravity","notes":"GA stable ID gemini-3.6-flash. Google reports fewer tool calls and 17% fewer output tokens than 3.5 Flash on the AA Index workload; Computer Use is preview. Artificial Analysis measured about 304 output tok/s at high thinking.","reasoning_capable":true,"effort_default":"medium","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":42,"reasoning":38,"knowledge":45,"comms":40,"multimodal":47,"agentic":42}},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","provider":"Google","tier":"legacy","released":"2026-05-19","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55.1,"swe_verified":78,"livecodebench":70,"aime":75,"tau2":68,"aaii_v4_1":50,"swe_bench_pro":55.1,"deep_swe_v1_1":37,"terminal_bench_2_1":76.2,"agents_last_exam":null,"quality_vs_fable":76.1,"api_in":1.5,"api_out":9,"api_cache_hit":0.375,"batch_discount":50,"tok_per_sec":165,"subscription":"Free CLI: 1000 req/day at 1M context","notes":"Shipped GA at Google I/O May 19, 2026. Still offered, but Gemini 3.6 Flash is the recommended migration target with stronger agentic results and lower output-token pricing.","reasoning_capable":true,"effort_default":"thinking-budget","effort_levels":["off","low","medium","high"],"cache_discount":0.25,"capability_levels":{"coding":30,"reasoning":30,"knowledge":40,"comms":30,"multimodal":40,"agentic":20}},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","tier":"fast","released":"2026-07-21","license":"proprietary","jurisdiction":"US","context":1048576,"context_label":"1M","swe_pro":54.2,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"aaii_v4_1":36,"swe_bench_pro":54.2,"deep_swe_v1_1":null,"terminal_bench_2_1":54,"agents_last_exam":null,"quality_vs_fable":65.9,"api_in":0.3,"api_out":2.5,"api_cache_hit":0.03,"batch_discount":50,"tok_per_sec":490,"subscription":"Gemini API · AI Studio · Gemini app rollout","notes":"GA stable ID gemini-3.5-flash-lite. Google positions it for high-volume subagents, document parsing and structured extraction; Artificial Analysis measured about 490 output tok/s. Computer Use availability differs across Google documentation and should be validated per API surface.","reasoning_capable":true,"effort_default":"minimal","effort_levels":["minimal","low","medium","high"],"cache_discount":0.1,"capability_levels":{"coding":32,"reasoning":28,"knowledge":34,"comms":32,"multimodal":38,"agentic":35}},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash Lite","provider":"Google","tier":"legacy","released":"2026-03-10","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":22,"swe_verified":65,"livecodebench":60,"aime":68,"tau2":58,"api_in":0.25,"api_out":1.5,"api_cache_hit":0.0625,"batch_discount":50,"tok_per_sec":220,"subscription":"Free CLI + AI Studio","notes":"Lite tier; 1M context retained.","reasoning_capable":true,"effort_default":"off","effort_levels":["off","low","medium"],"cache_discount":0.25,"capability_levels":{"coding":22,"reasoning":22,"knowledge":30,"comms":25,"multimodal":35,"agentic":15}},{"id":"deepseek-v4-pro","name":"DeepSeek V4-Pro","provider":"DeepSeek","tier":"frontier","released":"2026-04-24","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":55.4,"swe_verified":80.6,"livecodebench":93.5,"aime":88,"tau2":78,"api_in":0.435,"api_out":0.87,"api_cache_hit":0.043,"batch_discount":null,"tok_per_sec":60,"subscription":"API only","notes":"MIT-licensed. Permanent pricing May 22, 2026 — Opus-class quality at ~1/10 cost. 1.6T/49B MoE, 1M context.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.0083,"capability_levels":{"coding":47,"reasoning":40,"knowledge":38,"comms":30,"multimodal":8,"agentic":28}},{"id":"deepseek-v4-flash","name":"DeepSeek V4-Flash","provider":"DeepSeek","tier":"fast","released":"2026-07-31","license":"MIT","jurisdiction":"China","context":1000000,"context_label":"1M","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"tau2":72,"terminal_bench_2_1":82.7,"agents_last_exam":25.2,"api_in":0.14,"api_out":0.28,"api_cache_hit":0.0028,"batch_discount":null,"tok_per_sec":90,"subscription":"API + open weights","notes":"V4-Flash-0731 public beta; 284B/13B active MoE, 1M context and 384K max output. DeepSeek reports Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in its unreleased minimal harness; internal DSBench results are not normalized here. A future 2x peak-hours rate is announced, but no effective date is published.","reasoning_capable":true,"effort_default":"think","effort_levels":["non-think","think"],"cache_discount":0.02,"capability_levels":{"coding":40,"reasoning":30,"knowledge":30,"comms":25,"multimodal":5,"agentic":25}},{"id":"minimax-m2.7","name":"MiniMax M2.7","provider":"MiniMax","tier":"frontier","released":"2026-03-18","license":"open-weight","jurisdiction":"China","context":205000,"context_label":"205K","swe_pro":40,"swe_verified":74,"livecodebench":72,"aime":80,"tau2":73,"api_in":0.3,"api_out":1.2,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":70,"subscription":"API + open-weight","notes":"Current flagship reasoner; MoE 230B/10B active.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.52,"capability_levels":{"coding":30,"reasoning":30,"knowledge":30,"comms":30,"multimodal":20,"agentic":20}},{"id":"grok-4.3","name":"Grok 4.3","provider":"xAI","tier":"frontier","released":"2026-05-06","license":"proprietary","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":45,"swe_verified":80,"livecodebench":80,"aime":88,"tau2":76,"api_in":1.25,"api_out":2.5,"api_cache_hit":0.31,"batch_discount":null,"tok_per_sec":75,"subscription":"SuperGrok $30 · Heavy $300","notes":"Current xAI flagship — aggressive pricing for frontier tier. Hybrid reasoning. Multimodal text+image.","reasoning_capable":true,"effort_default":"reasoning-on","effort_levels":["off","on"],"cache_discount":0.25,"capability_levels":{"coding":42,"reasoning":44,"knowledge":42,"comms":34,"multimodal":34,"agentic":34}},{"id":"grok-4.1-fast","name":"Grok 4.1 Fast","provider":"xAI","tier":"fast","released":"2026-03-20","license":"proprietary","jurisdiction":"US","context":2000000,"context_label":"2M","swe_pro":30,"swe_verified":70,"livecodebench":65,"aime":75,"tau2":65,"api_in":0.2,"api_out":0.5,"api_cache_hit":0.05,"batch_discount":null,"tok_per_sec":140,"subscription":"SuperGrok $30","notes":"Cheapest large-context model on market. 2M context, $0.05/Mtok cached input.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":0.25,"capability_levels":{"coding":28,"reasoning":28,"knowledge":30,"comms":26,"multimodal":26,"agentic":22}},{"id":"mistral-medium-3.5","name":"Mistral Medium 3.5","provider":"Mistral","tier":"frontier","released":"2026-04-29","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":42,"swe_verified":77.6,"livecodebench":75,"aime":75,"tau2":65,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":75,"subscription":"Le Chat Pro €20","notes":"EU jurisdiction. 128B dense, Apache 2.0. Strongest non-Chinese open-weight coding agent. Vibe agents (GitHub/Linear/Jira/Sentry integrations). Pricing not verified; Medium 3 was $0.40/$2.00.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":40,"reasoning":30,"knowledge":40,"comms":30,"multimodal":10,"agentic":20}},{"id":"mistral-large-3","name":"Mistral Large 3","provider":"Mistral","tier":"frontier","released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","context":256000,"context_label":"256K","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"tau2":70,"api_in":2,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":65,"subscription":"API + Le Chat","notes":"Cheapest premium output price in market among Western frontier. 675B/41B MoE, Apache 2.0.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":38,"reasoning":36,"knowledge":40,"comms":36,"multimodal":20,"agentic":28}},{"id":"subq-1m-preview","name":"SubQ 1M-Preview","provider":"SubQ","tier":"frontier","released":"2026-05-15","license":"proprietary","jurisdiction":"US","context":12000000,"context_label":"12M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":1,"api_out":5,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":80,"subscription":"API preview","notes":"First commercial subquadratic (non-transformer) LLM. ~1/5 frontier cost on long-context tasks. Capability rating estimated — public benchmarks pending.","reasoning_capable":null,"effort_default":null,"effort_levels":[],"cache_discount":1,"capability_levels":{"coding":30,"reasoning":32,"knowledge":36,"comms":30,"multimodal":5,"agentic":28}},{"id":"nemotron-3-ultra","name":"Nvidia Nemotron 3 Ultra","provider":"Nvidia","tier":"frontier","released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"tau2":76,"api_in":1.5,"api_out":6,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":300,"subscription":"build.nvidia.com (closed API tier) + open weights","notes":"550B/55B MoE, hybrid Mamba-Transformer. ~300 tok/s. Open weights also available — see self_hosting. Capability levels are best estimates pending independent benchmarks.","reasoning_capable":true,"effort_default":"medium","effort_levels":["low","medium","high"],"cache_discount":1,"capability_levels":{"coding":42,"reasoning":42,"knowledge":42,"comms":34,"multimodal":10,"agentic":34}},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","provider":"Swiss AI Initiative","tier":"fast","released":"2026-06-15","license":"Apache 2.0","jurisdiction":"Switzerland","context":4096,"context_label":"4K","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Self-hosted weights","notes":"Largest newly released Apertus v1.1 distilled checkpoint. Dense 4.6B storage / 3.8B compute parameters, trained on 1.7T tokens, supports 1,811 languages, and ships in BF16 plus server and Apple-oriented quantizations. No comparable coding-agent, throughput, or API-price evidence is published.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null},{"id":"inkling","name":"Inkling","provider":"Thinking Machines Lab","tier":"frontier","released":"2026-07-15","license":"Apache 2.0","jurisdiction":"US","context":1000000,"context_label":"1M","swe_pro":null,"swe_verified":null,"livecodebench":null,"aime":null,"tau2":null,"api_in":null,"api_out":null,"api_cache_hit":null,"batch_discount":null,"tok_per_sec":null,"subscription":"Open weights (Hugging Face); fine-tuning and inference via the Tinker platform and third-party providers","notes":"975B-parameter multimodal MoE (41B active), 66-layer decoder-only transformer, 1M-token context. Text/image/audio input, text-only output. No first-party per-token API pricing published -- Thinking Machines monetizes via the Tinker fine-tuning platform rather than metered inference. A smaller Inkling-Small (12B active) companion model was released alongside it. Benchmark results are published on the model card across reasoning, agentic, coding, factuality, vision, audio, and safety categories but are not yet normalized into this roster's comparable benchmark set.","reasoning_capable":false,"effort_default":null,"effort_levels":[],"cache_discount":null}],"subscriptions":[{"provider":"Anthropic","tier":"Pro","price_usd":20,"limits":"~45 messages / 5h on Sonnet · limited Opus","models":"Sonnet 4.6, Haiku 4.5, limited Opus 4.8/4.7","features":"Chat only (Claude Code removed April 2026). From June 15, 2026 split into Chat pool + Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 5×","price_usd":100,"limits":"5× Pro quotas · ~225 msg/5h Sonnet · expanded Opus","models":"Full Opus 4.8/4.7 · Sonnet 4.6 · Haiku 4.5","features":"Cache reads included flat-rate. From June 15, 2026: Chat pool + separate Agent SDK credit pool."},{"provider":"Anthropic","tier":"Max 20×","price_usd":200,"limits":"20× Pro quotas","models":"All","features":"For heavy Opus users. From June 15, 2026: Chat + Agent SDK credit pools."},{"provider":"OpenAI","tier":"Plus","price_usd":20,"limits":"~80 GPT-5.4 msg/3h","models":"GPT-5.4 (limited) · GPT-5.4 Mini · o-series","features":"ChatGPT · GPTs · Codex CLI 30-150 tasks/5h"},{"provider":"OpenAI","tier":"Pro 5×","price_usd":100,"limits":"5× Plus quotas (new tier April 2026)","models":"GPT-5.4 Thinking unlimited","features":"Released as Anthropic Max competitor"},{"provider":"OpenAI","tier":"Pro","price_usd":200,"limits":"Effectively unlimited","models":"All including o3-pro","features":"Original premium tier"},{"provider":"Google","tier":"Gemini Advanced","price_usd":20,"limits":"Generous, soft caps","models":"Gemini 3.1 Pro · Gemini 3.6 Flash · Gemini 3.5 Flash-Lite","features":"Workspace integration · 1M context"},{"provider":"Google","tier":"Ultra","price_usd":100,"limits":"Higher quotas + Veo video","models":"All Gemini + research preview","features":"Veo 3 video · Project Mariner"},{"provider":"Mistral","tier":"Le Chat Pro","price_usd":22,"limits":"Generous","models":"Mistral Large 3 · Codestral","features":"EU jurisdiction"},{"provider":"xAI","tier":"SuperGrok","price_usd":30,"limits":"Generous","models":"Grok 4 · Grok 4 Heavy","features":"X integration"},{"provider":"DeepSeek","tier":"API only","price_usd":null,"limits":"Pay per token","models":"V4-Pro · V4-Flash","features":"Cheapest frontier API ($0.435/$0.87 permanent since May 22, 2026)"}],"agent_policies":[{"provider":"Anthropic","subscription_automated":"Prohibited","enforcement":"Active (OAuth blocked Apr 4)","first_party_exception":"`claude -p` pipe mode and Claude Code itself","api_required_for_automation":true,"cite":[35]},{"provider":"OpenAI","subscription_automated":"Prohibited (ToS)","enforcement":"Currently tolerated","first_party_exception":"Codex CLI uses subscription quota","api_required_for_automation":false},{"provider":"Google","subscription_automated":"Allowed in CLI","enforcement":"—","first_party_exception":"Gemini CLI free tier 1000 req/day","api_required_for_automation":false},{"provider":"Mistral","subscription_automated":"Allowed","enforcement":"—","first_party_exception":"—","api_required_for_automation":false}],"harnesses":[{"id":"claude-code","name":"Claude Code","vendor":"Anthropic","license":"proprietary","category":"CLI + IDE","stars":49000,"providers":["Anthropic"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":true,"computer_use":true,"lsp":true,"git":true,"memory":"CLAUDE.md","sandbox":"local","swe_pro":46,"pricing":"Subscription Pro/Max","sweet_spot":"SubagentStop hooks, /plugin list, requiredMinimumVersion managed setting, MCP fixes, Agent Teams.","stumbles":"Anthropic-only. Removed from standard Pro tier April 2026 — push to Max.","cite":[23,1]},{"id":"opencode","name":"OpenCode","vendor":"Anomaly","license":"MIT","category":"CLI + ACP","stars":188119,"providers":["75+: Anthropic*, OpenAI, Google, Mistral, Kimi, GLM, Ollama, LM Studio"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.18.4 (Jul 20, 2026): adaptive Kimi thinking controls, provider-defined reasoning options, restored Azure endpoints, and a rewritten desktop prompt input.","stumbles":"Anthropic OAuth blocked April 4 — must use API key.","cite":[24,35,72]},{"id":"codex-cli","name":"Codex CLI","vendor":"OpenAI","license":"Apache 2.0","category":"CLI + macOS app","stars":100232,"providers":["OpenAI"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"cloud","swe_pro":56.8,"pricing":"Plus $20 / Pro $200","sweet_spot":"v0.144.6 (Jul 18, 2026): refreshed GPT-5.6 Sol/Terra/Luna bundled instructions and corrected Codex context-window metadata to 272K.","stumbles":"OpenAI-only. No MCP, no hooks. Tightly coupled to apply_patch tool.","cite":[25,36,70]},{"id":"gemini-cli","name":"Gemini CLI","vendor":"Google","license":"Apache 2.0","category":"CLI","stars":106096,"providers":["Google"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"GEMINI.md","sandbox":"local","swe_pro":null,"pricing":"Free 1000 req/day","sweet_spot":"v0.51.0 (Jul 16, 2026): hardened sensitive-path and symlink handling, read-only macOS sandbox git config, and modern-model escape-sequence fixes.","stumbles":"Sunsetting to Antigravity CLI for free tier on June 18, 2026; paid Gemini/Enterprise keys retain access.","cite":[26,71]},{"id":"aider","name":"Aider","vendor":"paul-gauthier","license":"Apache 2.0","category":"CLI","stars":32000,"providers":["Anthropic*","OpenAI","Google","Ollama","100+"],"mcp":false,"skills":false,"hooks":false,"subagents":false,"voice":true,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"CONVENTIONS.md","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"Mature, lightweight, voice-input native. Pair-programmer mode. Active, last commit Mar 2026. 44k stars.","stumbles":"No MCP, no hooks. Less ambitious than newer harnesses.","cite":[27]},{"id":"cline","name":"Cline","vendor":"cline-bot","license":"Apache 2.0","category":"VS Code extension","stars":64886,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":false,"subagents":false,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".clinerules","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v4.0.10 (Jul 20, 2026): current release adds telemetry for consecutive-mistake-limit events; multi-editor and CLI surfaces remain available.","stumbles":"Anthropic OAuth blocked. Can be expensive on long sessions.","cite":[28,35,73]},{"id":"roo-code","name":"Roo Code","vendor":"RooVetGit","license":"Apache 2.0","category":"VS Code extension","stars":18000,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":true,"git":true,"memory":".roo","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v3.53.0 (Apr 23, 2026). Power-user Cline fork, model-agnostic, BYOK. Recent: GPT-5.5 via Codex, Opus 4.7 on Vertex, checkpoint nav.","stumbles":"Same OAuth situation. Configuration complexity.","cite":[29,35]},{"id":"cursor","name":"Cursor","vendor":"Anysphere","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".cursorrules","sandbox":"local","swe_pro":null,"pricing":"Hobby free · Pro $20 · Pro+ $60 · Ultra $200 · Teams $40/user","sweet_spot":"Cursor 3.5 (May 20, 2026): Cloud Agents (isolated VMs, multi-repo), Composer 2.5, Agents Window, parallel subagents.","stumbles":"Closed source. Lock-in. Subscription required for serious use.","cite":[31]},{"id":"windsurf","name":"Windsurf / Devin Desktop","vendor":"Cognition","license":"proprietary","category":"IDE","stars":null,"providers":["Anthropic","OpenAI","Google","custom"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"windsurfrules","sandbox":"local","swe_pro":null,"pricing":"Pro $20 · Max $200","sweet_spot":"Renamed to Devin Desktop, Agent Command Center kanban, embedded Devin cloud agent, SWE-1.6 model, multi-agent + worktrees.","stumbles":"Pro $20 (was $15); smaller ecosystem than Cursor.","cite":[32]},{"id":"goose","name":"Goose","vendor":"Block","license":"Apache 2.0","category":"Desktop + CLI","stars":51387,"providers":["Anthropic*","OpenAI","Google","Ollama","many"],"mcp":true,"skills":false,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":true,"lsp":false,"git":true,"memory":".goosehints","sandbox":"local","swe_pro":null,"pricing":"Free (BYOK)","sweet_spot":"v1.43.0 (Jul 14, 2026): per-message token/cost/TTFT/tok-s metrics, ACP reconnection, GPT-5.6 support, dynamic Ollama Cloud discovery, and expanded providers.","stumbles":"Anthropic OAuth blocked. Less mindshare than OpenCode.","cite":[30,35,74]},{"id":"omo","name":"OMO (Multi-model orchestrator)","vendor":"community","license":"MIT","category":"Multi-agent orchestrator","stars":54000,"providers":["multi-provider"],"mcp":true,"skills":true,"hooks":true,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":false,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Free","sweet_spot":"Rebrand from oh-my-opencode; multi-model orchestration. Wraps Claude Code, OpenCode, Codex, Kimi K2, DeepSeek V4, Gemini CLI.","stumbles":"Niche. Steep learning curve."},{"id":"hermes","name":"Hermes Agent","vendor":"Nous Research","license":"MIT","category":"Self-hosted multi-platform agent","stars":null,"providers":["OpenRouter-style multi-model"],"mcp":false,"skills":true,"hooks":false,"subagents":true,"voice":true,"remote":true,"computer_use":true,"lsp":false,"git":true,"memory":"persistent memory + auto-gen skills","sandbox":"local/docker/ssh/singularity/modal","swe_pro":null,"pricing":"Free (self-hosted)","sweet_spot":"v0.16.0. Persistent memory + auto-generated skills — learns your projects. Bridges Telegram/Discord/Slack/WhatsApp/Signal/Email/CLI. Natural-language cron for unattended runs. Parallel isolated subagents. Web search, browser automation, vision, image-gen, TTS.","stumbles":"Not coding-specialised — general autonomous agent. Setup requires self-host script. No first-party SWE benchmark.","cite":[]},{"id":"github-copilot-cli","name":"GitHub Copilot CLI","vendor":"GitHub","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["GitHub Models"],"mcp":false,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":true,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Bundled with Copilot Business/Enterprise","sweet_spot":"GA Feb 25, 2026. Specialized sub-agents (Explore, Task, Code Review, Plan), background delegation, autopilot.","stumbles":"Bundled-only — no standalone tier. GitHub-centric."},{"id":"amp-cli","name":"Amp CLI","vendor":"Sourcegraph","license":"proprietary","category":"CLI + IDE","stars":null,"providers":["Multi"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"AGENTS.md","sandbox":"local","swe_pro":null,"pricing":"Subscription (Sourcegraph)","sweet_spot":"Spun out as standalone company 2026. Runs as sidebar agent inside Zed via Terminal Threads.","stumbles":"Early standalone phase."},{"id":"zed","name":"Zed","vendor":"Zed Industries","license":"proprietary","category":"IDE","stars":null,"providers":["15 LLM providers + MCP"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":"—","sandbox":"local","swe_pro":null,"pricing":"Personal free (2k predictions) · Pro $10 · Business $30/seat","sweet_spot":"Rust-native editor with first-class agent panel + ACP host. Terminal Threads run Claude Code/Amp inline.","stumbles":"Editor first; agent layer still maturing."},{"id":"continue-dev","name":"Continue","vendor":"Continue","license":"Apache 2.0","category":"IDE extension + CLI","stars":null,"providers":["multi-provider"],"mcp":true,"skills":false,"hooks":false,"subagents":true,"voice":false,"remote":false,"computer_use":false,"lsp":true,"git":true,"memory":".continue","sandbox":"local","swe_pro":null,"pricing":"Solo $0 · Team/Company ~$10/dev/mo","sweet_spot":"Agent mode plan+execute. Continuous AI, Mission Control, shared PR/ticket workflows.","stumbles":"Newer agent features still stabilizing across providers."}],"self_hosting":{"hardware_options":[{"id":"vast-2xa6000","name":"Vast.ai 2× A6000","vram_gb":96,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Reference cloud-GPU setup for this dashboard. No upfront capex. Docker templates, SSH/Cloudflare Zero Trust. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-3xa6000","name":"Vast.ai 3× A6000","vram_gb":144,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Headroom for Qwen 3 235B-A22B Q6_K. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"vast-4xa6000","name":"Vast.ai 4× A6000","vram_gb":192,"cost_label":"Live marketplace price at provisioning","cost_per_hour":null,"type":"cloud","notes":"Llama 3.1 405B Q4 territory. Marketplace rates vary by host, region, reliability, storage, and rental type.","price_status":"dynamic_marketplace","price_checked":"2026-07-19","source":"https://docs.vast.ai/guides/instances/pricing"},{"id":"mbp-14-m4pro-64","name":"MacBook Pro 14\" M4 Pro 64GB","vram_gb":64,"cost_label":"~$3,200 capex","cost_per_hour":null,"type":"local","notes":"MoE sweet spot. 273 GB/s memory bandwidth."},{"id":"mbp-16-m5max-128","name":"MacBook Pro 16\" M5 Max 128GB","vram_gb":128,"cost_label":"~$5,500 capex","cost_per_hour":null,"type":"local","notes":"Best portable inference. ~545 GB/s bandwidth."},{"id":"mba-15-32","name":"MacBook Air 15\" 32GB","vram_gb":32,"cost_label":"~$1,900 capex","cost_per_hour":null,"type":"local","notes":"Hard 32GB ceiling. Limited to ~14B dense or 26B MoE Q4."},{"id":"contabo-2xl40s","name":"Contabo 2 x L40S","provider":"Contabo","vram_gb":96,"cost_label":"€1,502/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"EU-oriented fixed monthly configuration; 64 vCPU, 213 GB RAM, 3.5 TB storage, and 15 TB bandwidth are listed for the 2-GPU tier."},{"id":"contabo-1xh200","name":"Contabo 1 x H200 NVL","provider":"Contabo","vram_gb":141,"cost_label":"€2,149/mo fixed plan (excl. VAT)","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"published_monthly","price_checked":"2026-07-19","source":"https://contabo.com/en/gpu-cloud/","notes":"Fixed monthly large-memory option. Confirm location and availability before treating it as a sovereignty or latency fit."},{"id":"infomaniak-1xl40s","name":"Infomaniak 1 x L40S","provider":"Infomaniak","vram_gb":48,"cost_label":"Live calculator / availability validation","cost_per_hour":null,"type":"cloud","show_in_fit":false,"price_status":"calculator_or_request","price_checked":"2026-07-19","source":"https://www.infomaniak.com/en/hosting/public-cloud/prices","notes":"Swiss OpenStack option with dedicated GPU access and usage billing. Public pages list L40S availability but do not expose a stable crawlable SKU price; validate availability for the selected region."},{"id":"hyperstack-1xh200","name":"Hyperstack 1 x H200 SXM","provider":"Hyperstack","vram_gb":141,"cost_label":"$3.50/hr on demand","cost_per_hour":3.5,"type":"cloud","show_in_fit":false,"price_status":"published_on_demand","price_checked":"2026-07-19","source":"https://www.hyperstack.cloud/","notes":"Minute-accurate on-demand billing; reservation pricing starts at $2.45/hr. Validate region, storage, and availability."}],"models":[{"id":"gemma-4-26b-moe","name":"Gemma 4 26B-A4B MoE","params_total":26,"params_active":4,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":35,"livecodebench":77,"aime":88,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-3xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"vast-4xa6000":{"quant":"BF16","vram_used":52,"tok_per_sec":130},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":22,"tok_per_sec":75},"mbp-16-m5max-128":{"quant":"BF16","vram_used":52,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"MoE — only 4B active per token. Faster than dense models 5× its size."},{"id":"gemma-4-31b-dense","name":"Gemma 4 31B Dense","params_total":31,"params_active":31,"released":"2026-04-02","license":"Apache 2.0","jurisdiction":"US","swe_pro":38,"livecodebench":80,"aime":89,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-3xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"vast-4xa6000":{"quant":"BF16","vram_used":62,"tok_per_sec":35},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":19,"tok_per_sec":28},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":26,"tok_per_sec":38},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Highest quality open Gemma. Slower per-token (full 31B active)."},{"id":"qwen-3.6-plus","name":"Qwen 3.6 Plus","params_total":397,"params_active":17,"released":"2026-04-11","license":"Apache 2.0","jurisdiction":"China","swe_pro":50,"livecodebench":71.4,"aime":87,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":38},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":35},"vast-4xa6000":{"quant":"Q8_0","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1M context. Top open agentic coder. April 11 release.","swe_verified":68.2},{"id":"llama-4-maverick","name":"Llama 4 Maverick","params_total":400,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":42,"livecodebench":70,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":28},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":25},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q3_K_M","vram_used":88,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"400B / 17B active MoE. 1M context. Strong MMLU-Pro (80.5%) but coding behind Chinese labs.","swe_verified":72},{"id":"minimax-m2.5-open","name":"MiniMax M2.5 (open weights)","params_total":456,"params_active":46,"released":"2026-01-20","license":"open-weight","jurisdiction":"China","swe_pro":40,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":130,"tok_per_sec":30},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Eclipsed by DeepSeek V4 and Qwen 3.6 Plus. China jurisdiction. Maintained for niche workloads."},{"id":"deepseek-v4-pro-open","name":"DeepSeek V4-Pro (open weights)","params_total":1600,"params_active":49,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":55.4,"livecodebench":93.5,"aime":90,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q2_K","vram_used":188,"tok_per_sec":14},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"1.6T total / 49B active MoE. MIT license. Strongest open coder. Requires very large multi-GPU deployment."},{"id":"llama-4-scout","name":"Llama 4 Scout","params_total":109,"params_active":17,"released":"2026-04-05","license":"Llama 4 Community","jurisdiction":"US","swe_pro":36,"swe_verified":68,"livecodebench":66,"aime":78,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-3xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"vast-4xa6000":{"quant":"BF16","vram_used":88,"tok_per_sec":42},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":32,"tok_per_sec":22},"mbp-16-m5max-128":{"quant":"Q6_K","vram_used":45,"tok_per_sec":32},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"10M token context (longest in any open model). 109B / 17B active MoE."},{"id":"kimi-k2.6","name":"Kimi K2.6","params_total":235,"params_active":21,"released":"2026-05-01","license":"Modified MIT","jurisdiction":"China","swe_pro":47,"swe_verified":75,"livecodebench":78,"aime":84,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":36},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":32},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":18},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Best open-weight for sub-agent fan-out. Built for harness-driven parallel pipelines. Chinese-trained."},{"id":"glm-5.1","name":"GLM 5.1","params_total":358,"params_active":32,"released":"2026-04-22","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":76,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":30},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT license — rare among open frontier models besides DeepSeek. Strong for enterprise fine-tuning."},{"id":"mistral-small-4","name":"Mistral Small 4","params_total":24,"params_active":24,"released":"2026-04-18","license":"Apache 2.0","jurisdiction":"EU","swe_pro":28,"swe_verified":60,"livecodebench":58,"aime":68,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-3xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"vast-4xa6000":{"quant":"BF16","vram_used":48,"tok_per_sec":65},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":48,"tok_per_sec":38},"mbp-16-m5max-128":{"quant":"BF16","vram_used":48,"tok_per_sec":55},"mba-15-32":{"quant":"Q4_K_M","vram_used":14,"tok_per_sec":30}},"notes":"6.5B effective parameters. EU jurisdiction. Best on-device option."},{"id":"nemotron-3-ultra-550b-a55b-moe","name":"Nvidia Nemotron 3 Ultra","params_total":550,"params_active":55,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":55,"swe_verified":82,"livecodebench":80,"aime":86,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-4xa6000":{"quant":"Q3_K_M","vram_used":180,"tok_per_sec":25},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Hybrid Mamba-Transformer. 1M context. NVIDIA Open Model License. Computex June 4 launch."},{"id":"nemotron-3-nano-30b-a3b","name":"Nvidia Nemotron 3 Nano 30B-A3B","params_total":30,"params_active":3,"released":"2026-06-04","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":30,"swe_verified":70,"livecodebench":68,"aime":76,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":60,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q6_K","vram_used":24,"tok_per_sec":70},"mbp-16-m5max-128":{"quant":"BF16","vram_used":60,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":17,"tok_per_sec":50}},"notes":"Hybrid Mamba-Transformer Nano variant. Open weights."},{"id":"nemotron-nano-9b-v2","name":"Nvidia Nemotron Nano 9B v2","params_total":9,"params_active":9,"released":"2026-04-12","license":"NVIDIA Open Model License","jurisdiction":"US","swe_pro":22,"swe_verified":60,"livecodebench":55,"aime":65,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-3xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"vast-4xa6000":{"quant":"BF16","vram_used":18,"tok_per_sec":180},"mbp-14-m4pro-64":{"quant":"BF16","vram_used":18,"tok_per_sec":90},"mbp-16-m5max-128":{"quant":"BF16","vram_used":18,"tok_per_sec":120},"mba-15-32":{"quant":"Q4_K_M","vram_used":6,"tok_per_sec":65}},"notes":"Dense 9B. Strong instruction following at edge sizes."},{"id":"kimi-k2.6-1t","name":"Kimi K2.6 (1T MoE)","params_total":1000,"params_active":32,"released":"2026-04-20","license":"Modified MIT","jurisdiction":"China","swe_pro":58.6,"swe_verified":80.2,"livecodebench":82,"aime":88,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-3xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":34},"vast-4xa6000":{"quant":"Q6_K","vram_used":175,"tok_per_sec":28},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"Top open intelligence. 80.2% SWE-Verified, 58.6% SWE-Pro, AAII 54. 262K context. Modified MIT."},{"id":"glm-4.6","name":"GLM 4.6","params_total":355,"params_active":32,"released":"2025-09-15","license":"MIT","jurisdiction":"China","swe_pro":44,"swe_verified":73,"livecodebench":75,"aime":80,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":32},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":28},"vast-4xa6000":{"quant":"Q8_0","vram_used":168,"tok_per_sec":24},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":90,"tok_per_sec":16},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 200K context. Z.ai release."},{"id":"qwen3.6-35b-a3b","name":"Qwen 3.6 35B-A3B","params_total":35,"params_active":3,"released":"2026-04-16","license":"Apache 2.0","jurisdiction":"China","swe_pro":38,"swe_verified":74,"livecodebench":72,"aime":80,"fits":{"vast-2xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-3xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"vast-4xa6000":{"quant":"BF16","vram_used":70,"tok_per_sec":140},"mbp-14-m4pro-64":{"quant":"Q4_K_M","vram_used":22,"tok_per_sec":80},"mbp-16-m5max-128":{"quant":"BF16","vram_used":70,"tok_per_sec":110},"mba-15-32":{"quant":"Q4_K_M","vram_used":16,"tok_per_sec":55}},"notes":"Apache 2.0 MoE — strong $/quality for fast bulk inference."},{"id":"cohere-command-a-plus","name":"Cohere Command A+","params_total":218,"params_active":28,"released":"2026-05-22","license":"CC-BY-NC + commercial","jurisdiction":"Canada","swe_pro":42,"swe_verified":75,"livecodebench":70,"aime":78,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":30},"vast-3xa6000":{"quant":"Q6_K","vram_used":130,"tok_per_sec":26},"vast-4xa6000":{"quant":"Q8_0","vram_used":170,"tok_per_sec":22},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":92,"tok_per_sec":14},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"First open-weights Cohere release in &gt;1 year. Sparse-MoE multimodal."},{"id":"mistral-large-3-open","name":"Mistral Large 3 (open weights)","params_total":675,"params_active":41,"released":"2025-12-02","license":"Apache 2.0","jurisdiction":"EU","swe_pro":45,"swe_verified":78,"livecodebench":76,"aime":80,"fits":{"vast-2xa6000":{"quant":null,"vram_used":null,"tok_per_sec":null},"vast-3xa6000":{"quant":"Q3_K_M","vram_used":140,"tok_per_sec":22},"vast-4xa6000":{"quant":"Q4_K_M","vram_used":180,"tok_per_sec":20},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":null,"vram_used":null,"tok_per_sec":null},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"675B/41B active MoE. Apache 2.0. EU jurisdiction."},{"id":"deepseek-v4-flash-open","name":"DeepSeek V4-Flash (open weights)","params_total":284,"params_active":13,"released":"2026-04-24","license":"MIT","jurisdiction":"China","swe_pro":50,"swe_verified":79,"livecodebench":88,"aime":82,"fits":{"vast-2xa6000":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":45},"vast-3xa6000":{"quant":"Q6_K","vram_used":115,"tok_per_sec":40},"vast-4xa6000":{"quant":"Q8_0","vram_used":150,"tok_per_sec":34},"mbp-14-m4pro-64":{"quant":null,"vram_used":null,"tok_per_sec":null},"mbp-16-m5max-128":{"quant":"Q4_K_M","vram_used":78,"tok_per_sec":20},"mba-15-32":{"quant":null,"vram_used":null,"tok_per_sec":null}},"notes":"MIT, 1M context, 384K max output. Cheap large-context open option."}],"frameworks":[{"name":"llama.cpp","best_for":"Mac (MLX), broad GGUF support","notes":"Best Apple Silicon performance via Metal."},{"name":"vLLM","best_for":"Multi-GPU servers, throughput","notes":"Production serving. Tensor parallelism."},{"name":"Ollama","best_for":"Easiest setup, dev workflow","notes":"Wrapper around llama.cpp. One-line model pull."},{"name":"MLX","best_for":"Apple Silicon native","notes":"Apple's framework. Best M-series perf."}]},"strategy":{"current_recommendation":{"label":"Risk-tiered GPT-5.6 + Fable stack","monthly_usd":null,"components":["GPT-5.6 Luna low — high-volume subagents and routine transformations","Gemini 3.5 Flash-Lite minimal — throughput-first extraction, parsing and parallel subagents","Gemini 3.6 Flash medium — Google-first coding, computer-use and multimodal agent loops","GPT-5.6 Terra medium — default engineering, review and documentation","GPT-5.6 Sol high or Claude Fable 5 high — escalation for hard, long-horizon tasks","Qwen 3.6 35B-A3B or DeepSeek V4-Flash — local route for privacy-sensitive bulk work"],"rationale":"Route by task risk instead of one subscription. Terra and Luna now retain strong coding-agent quality at substantially lower list price; Fable remains the long-horizon option only where its 30-day retention requirement is acceptable. Monthly cost is workload-dependent and must be computed from measured token volume."},"alternatives":[{"label":"Dual subscription (Claude Max + OpenAI Pro)","monthly_usd":230,"rationale":"Adds GPT-5.4 SWE-bench Pro lead and Codex CLI cloud sandbox. Worth $100/mo only if you frequently hit hard issues where Sonnet 4.6 plateaus.","verdict":"Defer until you have measured Sonnet plateau frequency for 30 days."},{"label":"API-only (no subscriptions)","monthly_usd":200,"rationale":"Pure pay-per-use. Maximum flexibility. Loses the subscription cache advantage — same workload costs 1.5–5× more for power users.","verdict":"Worse economics for your usage volume. Skip."},{"label":"Self-hosted maximalist","monthly_usd":80,"rationale":"Vast.ai 24/7 with Gemma 4 + Qwen 3 + occasional API top-up for frontier-only tasks.","verdict":"Cheapest if quality plateau at ~88% of Opus is acceptable. Operational overhead is real."}],"routing":[{"tier":"Bulk (70%)","use_for":"Classification, simple edits, triage, log parsing","preferred":"Gemini 3.5 Flash-Lite minimal · GPT-5.6 Luna low · Qwen 3.6 35B-A3B when local","cost_label":"Gemini $0.30/$2.50; Luna $0.20/$1.20 Mtok before cache, batch and reasoning tokens"},{"tier":"Mid (25%)","use_for":"Multi-file edits, code review, refactors, docs","preferred":"GPT-5.6 Terra medium · Gemini 3.6 Flash medium for Google-first or multimodal work","cost_label":"Terra $2/$12; Gemini $1.50/$7.50 Mtok before cache, batch and reasoning tokens"},{"tier":"Premium (5%)","use_for":"Architecture, hard debugging, long-context refactors","preferred":"GPT-5.6 Sol high · Claude Fable 5 high for long-horizon autonomy","cost_label":"Sol $5/$30; Fable $10/$50 Mtok before reasoning tokens"}],"open_questions":["Will the Nvidia Nemotron Coalition (Mistral, Cursor, Black Forest Labs, Thinking Machines) actually deliver a frontier-open consortium, or fragment within a quarter?","How does the June 15 Anthropic billing restructure (Chat pool + Agent SDK credit pool) change Max economics for agentic workloads?","Does Gemini 3.6 Flash’s lower output price and stronger agentic performance displace 3.1 Pro for most coding workflows before Gemini 3.5 Pro arrives?","DeepSeek V4-Pro permanent pricing ($0.435/$0.87) — does it force closed providers (OpenAI, Anthropic) to cut headline prices?","SubQ 1M-Preview's subquadratic architecture — does the next year see other commercial non-transformer LLMs?"],"task_fit":{"description":"For each task type, one recommended pick per provider drawn from the full roster. Click any cell to see runners-up that the daily briefing also considered. Burn × is derived from the matrix and recomputes when you change the reference dropdown.","rows":[{"task":"Tiny edits / boilerplate","description":"Single-file syntax fixes, regex replacements, formatting, docstrings, import sorting","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fixed-rate, fastest paid Anthropic option"},"openai":{"model_id":"gpt-5.5","effort":"minimal","rationale":"Cheapest reasoning setting on flagship"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Fastest Google route for cheap, high-volume pattern work"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Effectively $0/tok on existing Vast.ai infra"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ minimal","gpt-5.4-nano @ minimal"],"anthropic":["sonnet-4.6 @ low"],"google":["gemini-3.1-pro @ off"],"self_hosted":["llama-3.1-70b @ Q6_K"]}},{"task":"Normal coding task","description":"Single-function implementation, straightforward bug fixes, simple feature additions","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best $/quality on Anthropic side; cache reads included in Max 5×"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-specialised, ~⅓ burn of GPT-5.5 for near-identical SWE-Pro"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"58.7% SWE-Pro and stronger agentic coding than 3.5 Flash at lower output price"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Highest-quality open Gemma at full BF16 in 96GB"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium","gpt-5.5 @ low"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ high"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Multi-file implementation","description":"Feature spanning 3-10 files, requires understanding cross-file context","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Best style/intent understanding for code spanning files"},"openai":{"model_id":"gpt-5.5","effort":"medium","rationale":"Strong on planning; medium gives consistent multi-file edits"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Improved multi-step coding loops, fewer unwanted edits and 1M context"},"self_hosted":{"model_id":"qwen-3.6-plus","effort":null,"rationale":"Largest open MoE; strong reasoning across files"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ high","gpt-5.4 @ high"],"anthropic":["opus-4.7 @ high"],"google":["gemini-3.1-pro @ high"],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Hard debugging / architecture","description":"Race conditions, performance bottlenecks, design decisions with long-term implications","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Default Claude Code effort for Opus; best long-context retention"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"SWE-Pro leader; high effort balances depth and burn"},"google":{"model_id":"gemini-3.1-pro","effort":"high","rationale":"Preview Pro remains the deepest Google reasoning route; compare 3.6 Flash before paying the premium"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open-weight models trail frontier ~10-15pts here; not yet ready for hardest cases"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh","gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ high (lower burn)"],"google":[],"self_hosted":["qwen-3-235b-a22b — viable for some cases"]}},{"task":"Very hard autonomous repo task","description":"Overnight runs, agentic loops, tasks you don't supervise turn-by-turn","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"max","rationale":"Autonomy earns max effort's premium; no human in the loop to correct"},"openai":{"model_id":"gpt-5.5","effort":"xhigh","rationale":"Highest available reasoning budget on OpenAI side"},"google":{"model_id":null,"effort":null,"rationale":"Gemini 3.6 Flash improves agent loops but still lacks a max-equivalent tier for the hardest unsupervised runs"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Not recommended — frontier models earn their cost on the hardest tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ xhigh — better but $100/mo gating"],"anthropic":["opus-4.7 @ xhigh — if max budget is too steep"],"google":[],"self_hosted":[]}},{"task":"Code review / PR comments","description":"Reviewing diffs, suggesting improvements, finding issues without writing code yourself","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"high","rationale":"Critical thinking at moderate burn; doesn't need Opus depth"},"openai":{"model_id":"gpt-5.3-codex","effort":"medium","rationale":"Coding-tuned reviewer at ⅓ burn of GPT-5.5"},"google":{"model_id":"gemini-3.6-flash","effort":"medium","rationale":"Better instruction following and fewer unwanted code edits than 3.5 Flash"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Critique tasks don't need frontier; quality plateau acceptable"}},"runner_up_per_provider":{"openai":["gpt-5.4 @ medium"],"anthropic":["opus-4.7 @ medium"],"google":["gemini-3.1-pro @ medium"],"self_hosted":["qwen-3-235b-a22b"]}},{"task":"Long-context refactor (50K+ tokens)","description":"Cross-cutting changes across a large codebase, where retention is the bottleneck","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"high","rationale":"1M context with retention that actually works; Opus's sweet spot"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"400K API context; degrades faster than Opus past 200K"},"google":{"model_id":"gemini-3.6-flash","effort":"high","rationale":"Google reports a 54% 1M-context MRCR score, roughly double 3.5 Flash"},"self_hosted":{"model_id":null,"effort":null,"rationale":"Open models cap around 200K useful context; not yet competitive"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ xhigh — if context fits in 256K"],"anthropic":["opus-4.7 @ xhigh"],"google":["gemini-3.1-pro @ high"],"self_hosted":[]}},{"task":"Subagent / delegated subtask","description":"Spawned by a main agent for focused work; quality threshold lower than user-facing","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Fast, cheap, no effort knob to manage"},"openai":{"model_id":"gpt-5.4-mini","effort":"medium","rationale":"Best subagent — 94% of GPT-5.4 coding at 6× less"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Purpose-built for high-volume subagents at roughly 490 output tok/s"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Zero per-token cost for high-volume subagent loops"}},"runner_up_per_provider":{"openai":["gpt-5.4-mini @ low","gpt-5.4-nano @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["llama-3.1-70b @ Q4_K_M"]}},{"task":"Documentation / explanation","description":"READMEs, code comments, tutorials, explaining existing code in prose","recommendations":{"anthropic":{"model_id":"sonnet-4.6","effort":"medium","rationale":"Best prose style on Anthropic side"},"openai":{"model_id":"gpt-5.4","effort":"medium","rationale":"Reasoning helps less here; medium-effort GPT-5.4 wins on $/quality"},"google":{"model_id":"gemini-3.6-flash","effort":"low","rationale":"Lower output price than 3.5 Flash with improved knowledge-work and document analysis"},"self_hosted":{"model_id":"gemma-4-31b-dense","effort":null,"rationale":"Open models are competitive for prose tasks"}},"runner_up_per_provider":{"openai":["gpt-5.5 @ low"],"anthropic":["haiku-4.5"],"google":["gemini-3.5-flash-lite @ low"],"self_hosted":["qwen-3-235b-a22b @ Q4_K_M"]}},{"task":"Test scaffolding / fixtures","description":"Boilerplate test files, fixture generation, mocking, parameterised test suites","recommendations":{"anthropic":{"model_id":"haiku-4.5","effort":"medium","rationale":"Pattern-following work; doesn't need reasoning depth"},"openai":{"model_id":"gpt-5.4-mini","effort":"low","rationale":"Pattern-heavy; minimal reasoning, fast turnaround"},"google":{"model_id":"gemini-3.5-flash-lite","effort":"minimal","rationale":"Lowest-latency current Google model for repetitive structured generation"},"self_hosted":{"model_id":"gemma-4-26b-moe","effort":null,"rationale":"Bulk test scaffolding is exactly where self-hosted earns out"}},"runner_up_per_provider":{"openai":["gpt-5.3-codex @ low"],"anthropic":["sonnet-4.6 @ low"],"google":[],"self_hosted":["gemma-4-31b-dense"]}},{"task":"Cybersecurity / vulnerability research","description":"Penetration testing, red-teaming, vulnerability discovery — tasks where cyber-specific safeguards apply","recommendations":{"anthropic":{"model_id":"opus-4.8","effort":"xhigh","rationale":"Most capable Anthropic model; cyber tasks require Cyber Verification Program"},"openai":{"model_id":"gpt-5.5","effort":"high","rationale":"Flagship reasoning; vetted security access via OpenAI"},"google":{"model_id":null,"effort":null,"rationale":"No cyber-specialized Gemini variant"},"self_hosted":{"model_id":"deepseek-v4-pro","effort":null,"rationale":"Open-weight without safety filtering; deploy in isolated environment"}},"runner_up_per_provider":{"openai":["gpt-5.5-pro @ high"],"anthropic":["opus-4.7 @ xhigh"],"google":[],"self_hosted":["qwen-3.6-plus","kimi-k2.6"]}}]}},"quota_burn_matrix":{"baseline_label":"selected reference model at medium effort = 1.00×","baseline_model_id":"fable-5","baseline_effort":"medium","methodology":"Burn ratios are displayed as multiples of the selected reference model at medium effort. Underlying values stored as absolute units anchored to gpt-5.5 medium = 1.00. Per-model ratios from API list pricing (firm, ±5%). Per-effort multipliers grounded in: Anthropic's published thinking-token budgets (low=skip, medium=~1k, high=5k, xhigh=10k, max=20k tokens); ArtificialAnalysis's measurement of Sonnet 4.6 max ≈ 3× Sonnet 4.5 on the Intelligence Index; nxcode.io / OpenAI guidance that xhigh ≈ 3-5× medium; ampcode's GPT-5.5 cost analysis. Per-effort multipliers are ±20%. NOTE: quality vs effort is not strictly monotonic per task — aggregate quality trends upward, but individual tasks can see high beat xhigh or medium beat high. OpenAI explicitly warns that 'high is not automatically better than medium'. Models that run in a separate preview quota bucket with no standard-pool burn are encoded as 0.00x and annotated until final token or credit rates are published.","stacking_multipliers":[{"name":"Fast mode (/fast on)","multiplier":2,"scope":"any cell"},{"name":"Cached input (long thread)","multiplier":0.6,"scope":"any cell"},{"name":"Plan mode (/plan)","multiplier":"uses high effort","scope":"overrides default"}],"openai_matrix":[{"model_id":"gpt-5.6-sol","minimal":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4,"note":"List-price blend equals the historical GPT-5.5 medium unit. Effort factors are planning estimates until measured token use is available."},{"model_id":"gpt-5.6-terra","minimal":0.18,"low":0.26,"medium":0.4,"high":0.68,"xhigh":1.08,"max":1.6,"note":"Two-fifths of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.6-luna","minimal":0.018,"low":0.026,"medium":0.04,"high":0.068,"xhigh":0.108,"max":0.16,"note":"Four percent of Sol's 30/70 list-price blend; effort factors remain estimates."},{"model_id":"gpt-5.5","minimal":0.2,"low":0.5,"medium":1,"high":2,"xhigh":3.5},{"model_id":"gpt-5.5-pro","minimal":null,"low":null,"medium":6,"high":12,"xhigh":21},{"model_id":"gpt-5.4","minimal":0.1,"low":0.25,"medium":0.5,"high":1,"xhigh":1.75},{"model_id":"gpt-5.4-mini","minimal":0.011,"low":0.027,"medium":0.053,"high":0.107,"xhigh":0.187},{"model_id":"gpt-5.4-nano","minimal":0.003,"low":0.007,"medium":0.013,"high":null,"xhigh":null},{"model_id":"gpt-5.3-codex","minimal":0.07,"low":0.17,"medium":0.33,"high":0.67,"xhigh":1.17},{"model_id":"gpt-5.2","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null},{"model_id":"gpt-5.2-codex","minimal":null,"low":0.14,"medium":0.27,"high":0.53,"xhigh":null}],"anthropic_matrix":[{"model_id":"fable-5","low":1.1,"medium":1.69,"high":2.87,"xhigh":4.56,"max":6.76,"note":"Always-on adaptive thinking. Values start from the $10/$50 list-price blend; effort multipliers are planning estimates rather than published token budgets."},{"model_id":"opus-4.8","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Current flagship. Same headline $5/$25. Fast Mode 3× cheaper ($10/$50) vs Opus 4.7."},{"model_id":"opus-4.7","low":0.56,"medium":1.13,"high":2.8,"xhigh":4.5,"max":7.9,"note":"Legacy as of May 28. Thinking tokens: low=skip, medium=~1k, high=5k, xhigh=10k, max=20k. New tokenizer +35% tokens vs 4.6."},{"model_id":"sonnet-4.6","low":0.25,"medium":0.5,"high":1.25,"xhigh":null,"max":3.5,"note":"API default high. No xhigh tier. AA measured max ≈ 3× Sonnet 4.5 cost."},{"model_id":"haiku-4.5","low":null,"medium":0.17,"high":null,"xhigh":null,"max":null,"note":"No effort control — fixed-rate model."}],"google_matrix":[{"model_id":"gemini-3.1-pro","off":0.13,"low":0.2,"medium":0.3,"high":0.5},{"model_id":"gemini-3.5-flash","off":0.02,"low":0.03,"medium":0.05,"high":0.09},{"model_id":"gemini-3.6-flash","off":0.017,"low":0.025,"medium":0.042,"high":0.076},{"model_id":"gemini-3.5-flash-lite","off":0.005,"low":0.008,"medium":0.014,"high":0.024}],"non_reasoning_note":"Models without reasoning controls: Haiku 4.5 (Anthropic matrix above as fixed rate), GPT-5.5 Instant (ChatGPT default since May 5), Mistral Medium 3.5, MiniMax M2.7, Kimi K2.6, GLM 4.6, Qwen 3.6 variants, all Llama 4 variants, Grok 4.1 Fast, SubQ 1M-Preview — burn at fixed rate per model regardless of effort knob. DeepSeek V4-Flash-0731 now exposes thinking and non-thinking modes.","effort_quality_factors":{"minimal":0.65,"low":0.85,"medium":0.94,"high":0.98,"xhigh":1,"max":1.02,"off":0.85,"on":0.95},"quality_methodology":"Quality % = (model_swe_pro / reference_swe_pro × effort_quality_factor) × 100. Effort quality factors: minimal=0.65, low=0.85, medium=0.94, high=0.98, xhigh=1.00, max=1.02. Anchors: Anthropic's Hex measurement (low Opus 4.7 ≈ medium Opus 4.6 quality) anchors low ≈ 0.85; apiyi.com's report that max gains ~3pts over xhigh on hardest tasks anchors max=1.02; ampcode's GPT-5.5 analysis (medium captures 'most' of capability) anchors medium=0.94. IMPORTANT: these are AGGREGATE estimates. stet.sh found per-task reversals — high can beat xhigh on some tasks, medium can beat high. Treat as ±10% indicators, not precise measures.","unit_anchor":"gpt-5.5 medium","unit_anchor_note":"All burn values are stored as absolute units anchored to gpt-5.5 medium = 1.00. At render time, each cell is divided by the reference model's medium-effort cell (or its single datapoint for non-reasoning models) to produce the displayed ×-of-reference number. When the reference dropdown changes, all burn ratios recalculate against the new reference.","workload_presets":{"cold":{"label":"Cold (0% cache)","description":"One-off prompts, no context reuse. Worst-case burn.","cache_hit_rate":0},"mixed":{"label":"Mixed (40% cache)","description":"Interactive coding with some context reuse. Typical default.","cache_hit_rate":0.4},"warm":{"label":"Warm (70% cache)","description":"Sustained agentic session. AA's industry-standard 7:2:1 blend assumes this rate. danielvaughan.com's Codex CLI model uses 70%.","cache_hit_rate":0.7},"hot":{"label":"Hot (90% cache)","description":"Long warmer-pattern sessions with stable system prompts and reused context. vsits.co documented 99% achievable on Claude Code Max subscription with optimization.","cache_hit_rate":0.9}},"default_workload":"mixed","cache_methodology":"Cache discount = (cache_hit_price / input_price). Lower = better discount. Effective burn = (1 - cache_hit_rate) × raw_burn + cache_hit_rate × raw_burn × cache_discount. Per-provider cache discounts: Anthropic 10% (90% off), OpenAI 25% on flagship/10% on 5.4 family, DeepSeek V4-Pro 0.83% (99.2% off — most aggressive in industry), DeepSeek V4-Flash 2%, MiniMax 52%, Mistral &amp; Grok no published cache pricing (modeled as 100% = no discount), Google ~25% plus per-hour storage fee (not modeled). IMPORTANT CAVEAT on Anthropic subscriptions: docs say cache reads count at 10% rate, but GitHub issue anthropics/claude-code#24147 reports cache reads burning quota at full rate on Max subscriptions. Treat the displayed cache benefit as accurate for API usage; subscription quota behavior is contested. Cache writes (1.25× / 2× of input rate on Anthropic) are not modeled — amortize across many reads in steady-state."},"sources":[{"n":1,"category":"Anthropic","title":"Anthropic news","url":"https://www.anthropic.com/news"},{"n":2,"category":"Anthropic","title":"Anthropic Claude model overview","url":"https://docs.claude.com/en/docs/about-claude/models/overview"},{"n":3,"category":"Anthropic","title":"Anthropic pricing","url":"https://www.anthropic.com/pricing"},{"n":4,"category":"OpenAI","title":"OpenAI news","url":"https://openai.com/news/"},{"n":5,"category":"OpenAI","title":"OpenAI model catalog","url":"https://platform.openai.com/docs/models"},{"n":6,"category":"OpenAI","title":"OpenAI API pricing","url":"https://openai.com/api/pricing/"},{"n":7,"category":"Google","title":"Google DeepMind blog","url":"https://blog.google/technology/google-deepmind/"},{"n":8,"category":"Google","title":"Gemini API models","url":"https://ai.google.dev/gemini-api/docs/models"},{"n":9,"category":"Google","title":"Gemini API pricing","url":"https://ai.google.dev/pricing"},{"n":10,"category":"Mistral","title":"Mistral news","url":"https://mistral.ai/news/"},{"n":11,"category":"Mistral","title":"Mistral La Plateforme pricing","url":"https://mistral.ai/products/la-plateforme#pricing"},{"n":12,"category":"xAI","title":"xAI blog (Grok)","url":"https://x.ai/blog"},{"n":13,"category":"DeepSeek","title":"DeepSeek API docs","url":"https://api-docs.deepseek.com/"},{"n":14,"category":"MiniMax","title":"MiniMax news","url":"https://www.minimaxi.com/en/news"},{"n":15,"category":"Benchmarks","title":"SWE-bench (Verified)","url":"https://www.swebench.com/","note":"SWE-bench Verified is widely considered contaminated; treat with skepticism."},{"n":16,"category":"Benchmarks","title":"SWE-bench Pro leaderboard","url":"https://scale.com/leaderboard/swe_bench_pro","note":"Preferred trustworthy code-agent benchmark."},{"n":17,"category":"Benchmarks","title":"LiveCodeBench","url":"https://livecodebench.github.io/"},{"n":18,"category":"Benchmarks","title":"Artificial Analysis (cross-provider pricing + benchmarks)","url":"https://artificialanalysis.ai/"},{"n":19,"category":"Benchmarks","title":"LMArena (chatbot arena)","url":"https://lmarena.ai/"},{"n":20,"category":"Open-weight","title":"Hugging Face — trending models","url":"https://huggingface.co/models?sort=trending"},{"n":21,"category":"Open-weight","title":"Hugging Face blog","url":"https://huggingface.co/blog"},{"n":22,"category":"Open-weight","title":"r/LocalLLaMA (community signal)","url":"https://www.reddit.com/r/LocalLLaMA/"},{"n":23,"category":"Harnesses","title":"Claude Code (anthropics/claude-code)","url":"https://github.com/anthropics/claude-code"},{"n":24,"category":"Harnesses","title":"OpenCode (sst/opencode)","url":"https://github.com/sst/opencode"},{"n":25,"category":"Harnesses","title":"Codex CLI (openai/codex)","url":"https://github.com/openai/codex"},{"n":26,"category":"Harnesses","title":"Gemini CLI (google-gemini/gemini-cli)","url":"https://github.com/google-gemini/gemini-cli"},{"n":27,"category":"Harnesses","title":"Aider (Aider-AI/aider)","url":"https://github.com/Aider-AI/aider"},{"n":28,"category":"Harnesses","title":"Cline (cline/cline)","url":"https://github.com/cline/cline"},{"n":29,"category":"Harnesses","title":"Roo Code (RooVetGit/Roo-Code)","url":"https://github.com/RooVetGit/Roo-Code"},{"n":30,"category":"Harnesses","title":"Goose (block/goose)","url":"https://github.com/block/goose"},{"n":31,"category":"Harnesses","title":"Cursor changelog","url":"https://cursor.com/changelog"},{"n":32,"category":"Harnesses","title":"Windsurf changelog","url":"https://windsurf.com/changelog"},{"n":33,"category":"Hardware","title":"Vast.ai (spot GPU pricing — A6000, H100)","url":"https://vast.ai/"},{"n":34,"category":"Hardware","title":"Apple MacBook Pro (M-series)","url":"https://www.apple.com/shop/buy-mac/macbook-pro"},{"n":35,"category":"Policy","title":"Anthropic OAuth third-party restrictions (Apr 4, 2026)","url":"https://www.anthropic.com/news"},{"n":36,"category":"Policy","title":"OpenAI Codex token-based pricing migration (Apr 2, 2026)","url":"https://openai.com/news/"},{"n":37,"category":"Aggregators","title":"Hacker News (AI tags)","url":"https://news.ycombinator.com/"},{"n":38,"category":"Benchmarks","title":"AIME (math competition benchmark)","url":"https://aimeproblems.com/","note":"Used to anchor Reasoning axis ratings."},{"n":39,"category":"Benchmarks","title":"MMLU / MMLU-Pro (knowledge benchmark)","url":"https://github.com/hendrycks/test","note":"Used to anchor Knowledge axis ratings."},{"n":40,"category":"Benchmarks","title":"tau-bench (tool-use + agentic behaviour)","url":"https://github.com/sierra-research/tau-bench","note":"Used to anchor Agentic axis ratings."},{"n":41,"category":"Nvidia","title":"build.nvidia.com","url":"https://build.nvidia.com/"},{"n":42,"category":"Nvidia","title":"Hugging Face — Nvidia","url":"https://huggingface.co/nvidia"},{"n":43,"category":"Nvidia","title":"Nvidia blogs","url":"https://blogs.nvidia.com/"},{"n":44,"category":"Benchmarks","title":"SWE-bench Pro public leaderboard (Scale)","url":"https://scale.com/leaderboard/swe_bench_pro_public"},{"n":45,"category":"Benchmarks","title":"Scale Labs","url":"https://labs.scale.com/"},{"n":46,"category":"Benchmarks","title":"SWE-Rebench","url":"https://swe-rebench.com/"},{"n":47,"category":"Benchmarks","title":"Terminal-Bench 2.0","url":"https://tbench.ai/leaderboard"},{"n":48,"category":"Aggregators","title":"OpenRouter","url":"https://openrouter.ai/"},{"n":49,"category":"Aggregators","title":"LLM-Stats","url":"https://llm-stats.com/"},{"n":50,"category":"Benchmarks","title":"Artificial Analysis Intelligence Index","url":"https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index"},{"n":51,"category":"Harnesses","title":"Claude Code changelog","url":"https://code.claude.com/docs/en/changelog"},{"n":52,"category":"Harnesses","title":"Cursor pricing","url":"https://cursor.com/pricing"},{"n":53,"category":"Harnesses","title":"Windsurf changelog (Devin Desktop)","url":"https://windsurf.com/changelog"},{"n":54,"category":"Harnesses","title":"Gemini CLI changelogs","url":"https://geminicli.com/docs/changelogs"},{"n":55,"category":"Harnesses","title":"Codex CLI releases","url":"https://github.com/openai/codex/releases"},{"n":56,"category":"Google","title":"DeepMind models","url":"https://deepmind.google/models"},{"n":57,"category":"Mistral","title":"Mistral pricing","url":"https://mistral.ai/pricing"},{"n":58,"category":"Google","title":"Gemini API pricing (canonical)","url":"https://ai.google.dev/gemini-api/docs/pricing"},{"n":59,"category":"MiniMax","title":"Artificial Analysis — MiniMax M2.7","url":"https://artificialanalysis.ai/models/minimax-m2-7"},{"n":60,"category":"Moonshot","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3"},{"n":61,"category":"Moonshot","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models"},{"n":62,"category":"OpenAI","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/"},{"n":63,"category":"Hosting","title":"Runpod GPU cloud pricing","url":"https://www.runpod.io/pricing"},{"n":64,"category":"Hosting","title":"Runpod July 2024 GPU price changes","url":"https://www.runpod.io/blog/runpod-slashes-gpu-prices-more-power-less-cost-for-ai-builders"},{"n":65,"category":"Hosting","title":"Contabo GPU Cloud configurations and pricing","url":"https://contabo.com/en/gpu-cloud/"},{"n":66,"category":"Hosting","title":"Infomaniak Public Cloud pricing","url":"https://www.infomaniak.com/en/hosting/public-cloud/prices"},{"n":67,"category":"Hosting","title":"Infomaniak GPU flavor documentation","url":"https://docs.infomaniak.cloud/compute/instances/flavors/"},{"n":68,"category":"Hosting","title":"Hyperstack GPU cloud pricing","url":"https://www.hyperstack.cloud/"},{"n":69,"category":"Hosting","title":"Vast.ai marketplace pricing methodology","url":"https://docs.vast.ai/guides/instances/pricing"},{"n":70,"category":"Harnesses","title":"Codex CLI 0.144.6 release","url":"https://github.com/openai/codex/releases/tag/rust-v0.144.6"},{"n":71,"category":"Harnesses","title":"Gemini CLI 0.51.0 release","url":"https://github.com/google-gemini/gemini-cli/releases/tag/v0.51.0"},{"n":72,"category":"Harnesses","title":"OpenCode 1.18.4 release","url":"https://github.com/anomalyco/opencode/releases/tag/v1.18.4"},{"n":73,"category":"Harnesses","title":"Cline 4.0.10 release","url":"https://github.com/cline/cline/releases/tag/v4.0.10"},{"n":74,"category":"Harnesses","title":"Goose 1.43.0 release","url":"https://github.com/aaif-goose/goose/releases/tag/v1.43.0"},{"n":75,"category":"Moonshot","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing"},{"n":76,"category":"Swiss AI Initiative","title":"Apertus v1.1 4B Instruct model card","url":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"n":77,"category":"Google","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"n":78,"category":"Google","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"n":79,"category":"Google","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"n":80,"category":"Google","title":"Gemini API release notes","url":"https://ai.google.dev/gemini-api/docs/changelog"},{"n":81,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash"},{"n":82,"category":"Benchmarks","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"},{"n":83,"category":"OpenAI","title":"GPT-5.6 launch and benchmark table","url":"https://openai.com/index/gpt-5-6/"},{"n":84,"category":"Google","title":"Gemini API Managed Agents update","url":"https://blog.google/innovation-and-ai/technology/developers-tools/expanding-managed-agents-gemini-api/"},{"n":85,"category":"Benchmarks","title":"Scale Labs SWE-bench Pro public leaderboard","url":"https://labs.scale.com/api/pdf/leaderboard/swe_bench_pro_public","note":"Public leaderboard values use their own model versions and evaluation setup; do not merge them mechanically with vendor launch tables."},{"n":86,"category":"Anthropic","title":"What's new in Claude Opus 5","url":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5","note":"Official launch specifications, pricing, availability, and migration behaviour."},{"n":87,"category":"Benchmarks","title":"Artificial Analysis — Claude Opus 5","url":"https://artificialanalysis.ai/models/claude-opus-5","note":"Independent effort-specific intelligence, latency, throughput, and price analysis."},{"n":88,"category":"Thinking Machines Lab","title":"Inkling model card","url":"https://thinkingmachines.ai/model-card/inkling/"},{"n":89,"category":"OpenAI","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra"},{"n":90,"category":"OpenAI","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna"},{"n":91,"category":"DeepSeek","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","note":"Vendor-reported benchmark results use DeepSeek Harness minimal mode at max effort; DSBench results are internal."},{"n":92,"category":"DeepSeek","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/"},{"n":93,"category":"OpenAI","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","note":"Official July 30 pricing, paid-subscription credit, and Sol API Fast mode details."},{"n":94,"category":"OpenAI","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","note":"Official August 6 ChatGPT product update; it does not change the API model contract."}],"provider_sources":{"Anthropic":[1,2,3,86,87],"OpenAI":[4,5,6,62,83,89,90,93,94],"Google":[7,8,9,58,77,78,79,80,81,82,84],"Mistral":[10,11],"xAI":[12],"DeepSeek":[13,91,92],"MiniMax":[14],"Meta":[20,21],"Alibaba":[20,21],"Moonshot":[20,21,60,61,75],"Zhipu":[20,21],"Nvidia":[41,42,43],"Cohere":[20,21],"Swiss AI Initiative":[76],"SubQ":[],"Thinking Machines Lab":[88]},"section_sources":{"models":[1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,60,61,62,76,77,78,79,80,81,82,86,87,89,90,91,92],"harnesses":[23,24,25,26,27,28,29,30,31,32,35,51,52,53,54,55,70,71,72,73,74],"self_hosting":[20,21,22,33,34,63,64,65,66,67,68,69,76],"strategy":[1,3,4,6,16,18,33,35,36]},"capabilities":{"schema_version":"1.0","axes":[{"key":"coding","label":"Coding","short":"Code","effort_sensitivity":0.5},{"key":"reasoning","label":"Reasoning &amp; Architecture","short":"R&amp;A","effort_sensitivity":0.6},{"key":"knowledge","label":"Knowledge &amp; Research","short":"K&amp;R","effort_sensitivity":0.05},{"key":"comms","label":"Communication &amp; Docs","short":"Comms","effort_sensitivity":0.2},{"key":"multimodal","label":"Multimodal","short":"MM","effort_sensitivity":0.05},{"key":"agentic","label":"Agentic","short":"Agent","effort_sensitivity":0.4}],"level_labels":{"coding":{"1":"Snippet","2":"Standard","3":"Cross-file","4":"Hard","5":"Frontier"},"reasoning":{"1":"Apply known","2":"Multi-step","3":"Cross-cutting","4":"Novel system","5":"Research-grade"},"knowledge":{"1":"Recall","2":"Contextual","3":"Cross-domain","4":"Frontier","5":"Original"},"comms":{"1":"Grammatical","2":"Structured","3":"Tutorial","4":"Editorial","5":"Publishable"},"multimodal":{"1":"Text only","2":"Image-in","3":"Image reasoning","4":"Video/audio","5":"Cross-modal gen"},"agentic":{"1":"Single-turn","2":"Tool calls","3":"Plan coherence","4":"Self-correcting","5":"Long-horizon"}},"focus_presets":{"balanced":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"coding-focused":{"coding":2,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"architecture-focused":{"coding":1,"reasoning":2,"knowledge":1,"comms":1,"multimodal":1,"agentic":1},"research-focused":{"coding":1,"reasoning":1,"knowledge":2,"comms":1,"multimodal":1,"agentic":1},"writing-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":2,"multimodal":1,"agentic":1},"multimodal-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":2,"agentic":1},"agentic-focused":{"coding":1,"reasoning":1,"knowledge":1,"comms":1,"multimodal":1,"agentic":2}},"default_focus":"balanced","methodology":"1-5 levels per axis; Coding rated against SWE-Pro / SWE-Verified evidence; Reasoning against AIME / LiveCodeBench / public hard-reasoning evals; Multimodal against published modality support (image/audio/video in/out); Agentic rated holding harness constant at Claude-Code-class baseline (plan coherence, tool-call quality, self-correction, calibrated stopping, refusal hygiene); Knowledge from model-card claims + community-validated state-of-art awareness; Comms from prose-evaluation rounds + structural quality. Ratings carry citation; see src/dashboard-context.md for the full rubric.","effort_formula":"effective(axis, effort) = max(0, ceiling[axis] × (1 − effort_sensitivity[axis] × (1 − radar_effort_factor[effort]))). radar_effort_factor[medium] = 1.00 → primary polygon equals reference polygon shape at medium effort for the same model. Lower efforts shrink the polygon by sensitivity-weighted amounts; higher efforts grow it and may push vertices OUTSIDE the outer level-50 ring — the ring is a rubric marker, not a hard cap. Vertices that exceed 50 are drawn in accent-hot to signal 'boosted above medium-effort baseline'.","radar_effort_factors":{"minimal":0.5,"low":0.75,"medium":1,"high":1.08,"xhigh":1.15,"max":1.22},"scale":{"rubric_min":1,"rubric_max":50,"visual_max":60,"band_size":10,"boost_band":"51–60","interpretation":"capability_levels[axis] = the model's effective capability at MEDIUM effort. Briefings rate from 1 to 50. The radar's visual scale extends to 60 to accommodate the 'effort-boost band' (51–60) — values that fall there are formula-derived from extra reasoning effort, never directly rated. The outer rubric ring at 50 is visually emphasized as 'current frontier'.","bands":{"1-10":"Snippet","11-20":"Standard","21-30":"Cross-file","31-40":"Hard","41-50":"Frontier","51-60":"Effort-boost band (derived, not rated)"}}},"report_metrics":{"schema_version":"1.0","researched_at":"2026-08-01","currency":"USD","reference_options":["fable-5","gpt-5.6-sol","gpt-5.6-terra","gpt-5.6-luna","opus-4.8","gpt-5.5","gpt-5.5-pro","gpt-5.5-instant","kimi-k3","gemini-3.6-flash","gemini-3.5-flash-lite"],"default_reference":"fable-5","methodology":{"benchmark_scale":"Each benchmark is divided by benchmark.max_value and expressed on a 0-100 scale. The quality composite is the weighted arithmetic mean of available normalized benchmarks; it is shown only when coverage is at least 0.50.","quality_weights":{"aaii_v4_1":0.2,"coding_agent_index_v1_1":0.2,"swe_bench_pro":0.2,"deep_swe_v1_1":0.15,"terminal_bench_2_1":0.15,"agents_last_exam":0.1},"speed_score":"100 * sqrt(task_speed_index / max_task_speed_index). task_speed_index is 100 * reference_time / model_time. It is not output tokens per second and is comparable only within the cited evaluation setup.","cost_score":"10 + 90 * ln(max_blended_price / blended_price) / ln(max_blended_price / min_blended_price). Blended price is 0.30 * input_price + 0.70 * output_price per million tokens. Higher cost score is better.","scq_compound":"Geometric mean of quality_score, speed_score and cost_score. Geometric mean prevents one category from fully compensating for a weak category. Null when quality coverage is below 0.50.","capability_compound":"Weighted arithmetic mean of the six capability axes. Balanced uses weight 1 for each axis; a selected focus uses weight 2.5 for that axis and 1 for every other axis.","reference_quality":"Each model stores or derives a Fable-anchored quality value. Displayed quality = 100 * model_anchor / selected_reference_anchor, so the selected reference is always exactly 100. New-suite composites use only explicitly overlapping evaluations; legacy rows fall back to SWE-Bench Pro. Unknown evidence remains null, never zero.","burn":"Raw blended-price units are multiplied by an effort factor and cache factor, then divided by the selected reference model at medium effort. Cache factor = (1-hit_rate) + hit_rate*cache_read_ratio.","missing_data":"Keep unknown values null. Display insufficient comparable evidence rather than zero. Detailed benchmark cells remain version-specific and may be empty even when a reference-relative quality anchor exists from a separate documented comparison set."},"market_signals":[{"id":"frontier_quality","label":"Frontier quality index","explanation":"Best broad intelligence score published on Artificial Analysis Intelligence Index v4.1. Higher is better; the benchmark version must remain fixed across the series.","unit":"index","direction":"up","current":59.9,"history":[{"date":"2026-01-01","value":48.2},{"date":"2026-02-01","value":50.1},{"date":"2026-03-01","value":52.7},{"date":"2026-04-01","value":55.7},{"date":"2026-05-01","value":55.7},{"date":"2026-06-01","value":59.9},{"date":"2026-07-01","value":59.9}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"coding_quality","label":"Coding-agent frontier","explanation":"Best score on Artificial Analysis Coding Agent Index v1.1. Higher means better end-to-end coding-agent performance, not just code completion.","unit":"index","direction":"up","current":80,"history":[{"date":"2026-01-01","value":66.1},{"date":"2026-02-01","value":69.8},{"date":"2026-03-01","value":72.5},{"date":"2026-04-01","value":76.4},{"date":"2026-05-01","value":77.2},{"date":"2026-06-01","value":77.2},{"date":"2026-07-01","value":80}],"source":"https://openai.com/index/gpt-5-6/"},{"id":"task_speed","label":"Coding task speed increase","explanation":"Fastest current end-to-end coding task-rate index, where Fable 5 is fixed at 100. A value of 320 means an estimated 3.2 times as many comparable tasks per unit time; it is not tokens per second.","unit":"Fable=100","direction":"up","current":320,"history":[{"date":"2026-01-01","value":100},{"date":"2026-02-01","value":112},{"date":"2026-03-01","value":126},{"date":"2026-04-01","value":145},{"date":"2026-05-01","value":180},{"date":"2026-06-01","value":180},{"date":"2026-07-01","value":320}],"source":"https://openai.com/index/gpt-5-6/","history_note":"Pre-July points are frozen planning estimates from prior report snapshots and should not be interpreted as one controlled longitudinal benchmark."},{"id":"blended_frontier_price","label":"Lowest frontier blended API price","explanation":"Lowest 30% input / 70% output list-price blend among models meeting the report's frontier quality floor. Lower is better; batch and caching are excluded.","unit":"$/Mtok","direction":"down","current":0.9,"history":[{"date":"2026-01-01","value":18},{"date":"2026-02-01","value":16.5},{"date":"2026-03-01","value":14},{"date":"2026-04-01","value":11.4},{"date":"2026-05-01","value":9},{"date":"2026-06-01","value":9},{"date":"2026-07-01","value":4.5},{"date":"2026-08-01","value":0.9}],"source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"open_weight_quality","label":"Open-weight quality vs selected reference","explanation":"Best open-weight model relative to the selected reference. Stored history is Fable-anchored and rebased in the browser whenever the reference changes. July uses Kimi K3's geometric mean across 14 overlapping launch-suite values; vendor-reported values require independent reproduction.","unit":"%","direction":"up","current":99.2,"history":[{"date":"2026-01-01","value":82},{"date":"2026-02-01","value":85},{"date":"2026-03-01","value":88},{"date":"2026-04-01","value":91},{"date":"2026-05-01","value":94},{"date":"2026-06-01","value":96},{"date":"2026-07-01","value":99.2}],"history_note":"Earlier points are frozen report estimates; July uses Moonshot’s K3 launch suite and is not a single controlled longitudinal benchmark.","source":"https://www.kimi.com/blog/kimi-k3"},{"id":"output_throughput","label":"Documented API output throughput","explanation":"Fastest provider-documented general API output rate in the tracked set. Unit is generated tokens per second; it is distinct from time to first token and end-to-end task speed.","unit":"tok/s","direction":"up","current":490,"history":[{"date":"2026-01-01","value":110},{"date":"2026-02-01","value":130},{"date":"2026-03-01","value":140},{"date":"2026-04-01","value":180},{"date":"2026-05-01","value":200},{"date":"2026-06-01","value":220},{"date":"2026-07-01","value":490}],"history_note":"Provider and independent figures use different serving and context conditions. The latest point is Artificial Analysis's Gemini 3.5 Flash-Lite measurement on Google's first-party API.","source":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite"}],"benchmarks":[{"id":"aaii_v4_1","label":"AA Intelligence Index v4.1","max_value":59.9,"unit":"index","good_direction":"high"},{"id":"coding_agent_index_v1_1","label":"AA Coding Agent Index v1.1","max_value":80,"unit":"index","good_direction":"high"},{"id":"swe_bench_pro","label":"SWE-Bench Pro","max_value":80,"unit":"%","good_direction":"high"},{"id":"deep_swe_v1_1","label":"DeepSWE v1.1","max_value":72.7,"unit":"%","good_direction":"high"},{"id":"terminal_bench_2_1","label":"Terminal-Bench 2.1","max_value":91.9,"unit":"%","good_direction":"high","note":"Maximum is GPT-5.6 Sol Ultra; ordinary Sol scores 88.8."},{"id":"agents_last_exam","label":"Agents' Last Exam","max_value":52.7,"unit":"%","good_direction":"high"}],"model_metrics":[{"model_id":"fable-5","name":"Claude Fable 5","provider":"Anthropic","benchmarks":{"aaii_v4_1":59.9,"coding_agent_index_v1_1":77.2,"swe_bench_pro":80,"deep_swe_v1_1":69.7,"terminal_bench_2_1":83.1,"agents_last_exam":40.5},"quality_score":94.9,"quality_coverage":1,"task_speed_index":100,"speed_score":50,"api_input":10,"api_cached_input":1,"api_output":50,"blended_price":38,"cost_score":10,"scq_compound":36.2,"speed_evidence":"Reference index; Anthropic labels comparative latency slower.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gpt-5.6-sol","name":"GPT-5.6 Sol","provider":"OpenAI","benchmarks":{"aaii_v4_1":58.9,"coding_agent_index_v1_1":80,"swe_bench_pro":64.6,"deep_swe_v1_1":72.7,"terminal_bench_2_1":88.8,"agents_last_exam":52.7},"quality_score":95.4,"quality_coverage":1,"task_speed_index":256,"speed_score":80,"api_input":5,"api_cached_input":0.5,"api_output":30,"blended_price":22.5,"cost_score":22.6,"scq_compound":55.7,"speed_evidence":"OpenAI introduced API Fast mode on July 30, claiming up to 2.5× Standard speed at 2× price with no intelligence change; the base task-speed index remains a separate vendor comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-terra","name":"GPT-5.6 Terra","provider":"OpenAI","benchmarks":{"aaii_v4_1":55,"coding_agent_index_v1_1":77.4,"swe_bench_pro":63.4,"deep_swe_v1_1":69.6,"terminal_bench_2_1":87.4,"agents_last_exam":50.4},"quality_score":91.8,"quality_coverage":1,"task_speed_index":300,"speed_score":86.6,"api_input":2,"api_cached_input":0.2,"api_output":12,"blended_price":9,"cost_score":44.6,"scq_compound":70.8,"speed_evidence":"OpenAI reports roughly one-third of Fable 5 task time for the family coding comparison.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"gpt-5.6-luna","name":"GPT-5.6 Luna","provider":"OpenAI","benchmarks":{"aaii_v4_1":51.2,"coding_agent_index_v1_1":74.6,"swe_bench_pro":62.7,"deep_swe_v1_1":67.2,"terminal_bench_2_1":84.7,"agents_last_exam":50.3},"quality_score":88.7,"quality_coverage":1,"task_speed_index":320,"speed_score":89.4,"api_input":0.2,"api_cached_input":0.02,"api_output":1.2,"blended_price":0.9,"cost_score":100,"scq_compound":92.6,"speed_evidence":"OpenAI calls Luna the fastest tier; 320 is a conservative planning index above the cited 300 family comparison and must not be treated as measured tok/s.","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"model_id":"kimi-k3","name":"Kimi K3","provider":"Moonshot AI","benchmarks":{"deep_swe_v1_1":67.5,"terminal_bench_2_1":88.3},"quality_score":null,"quality_coverage":0.33,"quality_vs_fable":99.2,"task_speed_index":null,"speed_score":null,"api_input":3,"api_cached_input":0.3,"api_output":15,"blended_price":11.4,"cost_score":38.9,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time or output-throughput figure published for K3 at launch.","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"model_id":"opus-4.8","name":"Claude Opus 4.8","provider":"Anthropic","benchmarks":{"aaii_v4_1":55.7,"coding_agent_index_v1_1":72.5,"swe_bench_pro":69.2,"deep_swe_v1_1":59,"terminal_bench_2_1":78.9,"agents_last_exam":45.2},"quality_score":87.7,"quality_coverage":1,"task_speed_index":180,"speed_score":67.1,"api_input":5,"api_cached_input":0.5,"api_output":25,"blended_price":19,"cost_score":26.7,"scq_compound":54,"speed_evidence":"Planning index based on moderate vendor latency and optional fast mode; validate in the local harness.","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"model_id":"gemini-3.1-pro","name":"Gemini 3.1 Pro Preview","provider":"Google","benchmarks":{"aaii_v4_1":46.5,"coding_agent_index_v1_1":42.7,"swe_bench_pro":54.2,"deep_swe_v1_1":11.8,"terminal_bench_2_1":70.7,"agents_last_exam":32.1},"quality_score":59.8,"quality_coverage":1,"task_speed_index":null,"speed_score":null,"api_input":null,"api_cached_input":null,"api_output":null,"blended_price":null,"cost_score":null,"scq_compound":null,"speed_evidence":"No comparable end-to-end task-time measurement in the selected source set.","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro"},{"model_id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","provider":"Google","benchmarks":{"aaii_v4_1":50,"swe_bench_pro":58.7,"deep_swe_v1_1":49,"terminal_bench_2_1":78},"quality_score":77.4,"quality_coverage":0.7,"quality_vs_fable":81.6,"task_speed_index":null,"speed_score":null,"api_input":1.5,"api_cached_input":0.15,"api_output":7.5,"blended_price":5.7,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 304 output tok/s at high thinking; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"model_id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","provider":"Google","benchmarks":{"aaii_v4_1":36,"swe_bench_pro":54.2,"terminal_bench_2_1":54},"quality_score":62.5,"quality_coverage":0.55,"quality_vs_fable":65.9,"task_speed_index":null,"speed_score":null,"api_input":0.3,"api_cached_input":0.03,"api_output":2.5,"blended_price":1.84,"cost_score":null,"scq_compound":null,"speed_evidence":"Artificial Analysis measured about 490 output tok/s; this is throughput, not comparable end-to-end task time.","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"}],"visualizations":{"bubble":{"x":"cost_score","y":"speed_score","size":"quality_vs_selected_reference","color":"provider","include_when":"quality vs reference, speed score and cost score are all non-null"},"heatmap":{"rows":"model","columns":["quality_vs_selected_reference","speed_score","cost_score"],"color_scale":"sequential_good","domain":[0,100]},"bcg":{"x":"capability_compound","y":"mean(quality_vs_selected_reference, speed_score, cost_score)","color":"provider","quadrants":"medians of currently visible models"}},"capability":{"axes":["coding","reasoning_architecture","knowledge_research","communication_docs","multimodal","agentic"],"focus_multiplier":2.5,"models":[{"model_id":"fable-5","scores":[98,99,96,96,90,99],"capability_compound":96.3,"scq_compound":36.2},{"model_id":"gpt-5.6-sol","scores":[98,97,95,95,96,98],"capability_compound":96.5,"scq_compound":55.7},{"model_id":"gpt-5.6-terra","scores":[95,92,91,92,90,95],"capability_compound":92.5,"scq_compound":70.8},{"model_id":"gpt-5.6-luna","scores":[92,88,86,89,86,91],"capability_compound":88.7,"scq_compound":92.6},{"model_id":"opus-4.8","scores":[92,94,94,95,82,94],"capability_compound":91.8,"scq_compound":54},{"model_id":"gemini-3.1-pro","scores":[78,84,94,84,98,72],"capability_compound":85,"scq_compound":null}],"provenance_note":"Capability scores are editorial rubric ratings synthesized from benchmark and feature evidence, not vendor benchmark results. Keep them separate from benchmark values."},"economics":{"effort_factors":{"none":0.45,"low":0.65,"medium":1,"high":1.7,"xhigh":2.7,"max":4},"cache_hit_presets":{"cold":0,"mixed":0.4,"warm":0.7,"hot":0.9},"models":[{"model_id":"fable-5","base_blended_price":38,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"],"note":"Always-on adaptive thinking; effort is a behavioral control, not a published token multiplier."},{"model_id":"gpt-5.6-sol","base_blended_price":22.5,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-terra","base_blended_price":9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"gpt-5.6-luna","base_blended_price":0.9,"cache_read_ratio":0.1,"efforts":["none","low","medium","high","xhigh","max"]},{"model_id":"opus-4.8","base_blended_price":19,"cache_read_ratio":0.1,"efforts":["low","medium","high","xhigh","max"]},{"model_id":"kimi-k3","base_blended_price":11.4,"cache_read_ratio":0.1,"efforts":["max"],"note":"Always-on reasoning at launch; the max-only effort control is behavioral and not a published token multiplier."}],"efficiency_filters":{"quality_floor_default":75,"provider_default":"all","effort_default":"medium","workload_default":"mixed","sort_default":"scq_compound_desc"}},"hardware_options":[{"id":"local-64gb","name":"64 GB unified-memory workstation","memory_gb":64,"type":"local","best_for":"3B-35B dense or small MoE models at Q4-Q8","cost_note":"Capex varies; compare measured memory bandwidth, not product year."},{"id":"local-128gb","name":"128 GB unified-memory workstation","memory_gb":128,"type":"local","best_for":"Up to roughly 100 GB quantized weights with headroom for KV cache","cost_note":"Portable and quiet; slower than datacenter GPUs for sustained batches."},{"id":"cloud-2xa6000","name":"2 x RTX A6000 48 GB","memory_gb":96,"type":"cloud","best_for":"70B-class dense and mid-size MoE Q4 deployments","cost_note":"Spot price varies by host; record price and interconnect at test time."},{"id":"cloud-h200","name":"1 x H200 141 GB","memory_gb":141,"type":"cloud","best_for":"High-throughput 70B inference and larger quantized MoE models","cost_note":"Use provider quote; hourly rates change frequently."},{"id":"cloud-b200","name":"1 x B200 192 GB","memory_gb":192,"type":"cloud","best_for":"Large-model throughput where software stack supports Blackwell","cost_note":"Availability and hourly rates vary; validate framework support."}],"hardware_model_fit":[{"model_id":"qwen3.6-35b-a3b","hardware_id":"local-64gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":55,"evidence":"planning_estimate"},{"model_id":"qwen3.6-35b-a3b","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"BF16","estimated_output_tps":140,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"local-64gb","fit_status":"unsupported","reason":"Quantized weights exceed usable memory."},{"model_id":"kimi-k2.6-1t","hardware_id":"local-128gb","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":16,"evidence":"planning_estimate"},{"model_id":"kimi-k2.6-1t","hardware_id":"cloud-h200","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":36,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-flash-open","hardware_id":"cloud-2xa6000","fit_status":"supported","quant":"Q4_K_M","estimated_output_tps":45,"evidence":"planning_estimate"},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-h200","fit_status":"unsupported","reason":"Single-device memory is insufficient at the tracked quantization."},{"model_id":"deepseek-v4-pro-open","hardware_id":"cloud-b200","fit_status":"supported","quant":"Q2_K","estimated_output_tps":14,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"local-64gb","fit_status":"supported","quant":"BF16","estimated_output_tps":38,"evidence":"planning_estimate"},{"model_id":"mistral-small-4","hardware_id":"cloud-h200","fit_status":"supported","quant":"BF16","estimated_output_tps":80,"evidence":"planning_estimate"}],"decision_support":{"recommended_stack":[{"route":"hardest_long_horizon","primary":"fable-5","effort":"high","fallback":"gpt-5.6-sol","why":"Use Fable where long-horizon reliability warrants cost and retention policy is acceptable."},{"route":"default_engineering","primary":"gpt-5.6-terra","effort":"medium","fallback":"gpt-5.6-sol","why":"Best balanced default in the current benchmark/cost set."},{"route":"high_volume_subagents","primary":"gpt-5.6-luna","effort":"low","fallback":"gemini-3.5-flash-lite","why":"Luna retains stronger coding-agent evidence; Flash-Lite is the throughput-first fallback for extraction and delegated subtasks."},{"route":"privacy_local","primary":"qwen3.6-35b-a3b","effort":null,"fallback":"deepseek-v4-flash-open","why":"Local-first route where external retention is unacceptable."},{"route":"vision_long_context","primary":"gemini-3.1-pro","effort":"high","fallback":"gpt-5.6-sol","why":"Prefer for multimodal context; validate preview stability before production."}],"routing_rules":["Route by task risk and evidence, not provider family.","Start bulk work on Luna or Terra and escalate only after an explicit verification failure.","Do not send ZDR-required data to Fable 5 because the model requires 30-day retention.","Record model, effort, cache state, region and harness version for every internal speed comparison.","Use self-hosted routes only when the selected hardware row is supported; never infer performance from an empty cell."]},"actions":[{"priority":"P0","owner":"Executive sponsor — define three transformation outcomes with measurable business and engineering baselines; avoid scaling pilots that have no accountable owner or adoption target.","status":"open","order":1},{"priority":"P0","owner":"Technology leadership — establish a model portfolio policy with capability, data-classification, regional, fallback, and retirement rules instead of standardizing on one provider.","status":"open","order":2},{"priority":"P0","owner":"Platform and finance — instrument end-to-end quality, latency, retries, human rework, and cost for representative workflows before negotiating capacity or subscriptions.","status":"open","order":3},{"priority":"P1","owner":"Security and legal — approve reusable controls for retention, training use, tool permissions, audit evidence, and human escalation by data class.","status":"open","order":4},{"priority":"P1","owner":"Engineering leadership — run a 30-task quarterly evaluation across one frontier, one balanced, one fast, and one open-weight route using identical harness conditions.","status":"open","order":5},{"priority":"P2","owner":"Infrastructure — select one sovereignty or resilience workload for an open-weight pilot and publish its full hardware, utilization, staffing, and throughput economics.","status":"open","order":6}],"changelog":[{"date":"2026-08-07","tag":"pricing","text":"Corrected the price-change history using OpenAI's July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. Added the paid Codex and ChatGPT Work credit reduction, unchanged subscription prices and quota budgets, and Sol API Fast mode (up to 2.5× Standard speed at 2× price)."},{"date":"2026-08-01","tag":"pricing","text":"Refreshed OpenAI's live GPT-5.6 API rate card: Terra is now $2/$12 per million input/output tokens and Luna is $0.20/$1.20, with cached input at $0.20 and $0.02 respectively. Recomputed the 30/70 workload blend, cost-efficiency scores, SCQ compounds and burn baselines. Superseded on August 7 with the official July 30 effective date and full change details."},{"date":"2026-08-01","tag":"benchmark","text":"Recorded DeepSeek V4-Flash-0731's July 31 public-beta evidence: Terminal-Bench 2.1 82.7 and Agents' Last Exam 25.2 at max effort in DeepSeek Harness minimal mode. Additional vendor-reported NL2Repo, Cybergym, Toolathlon and Automation Bench results remain narrative evidence rather than new shared columns."},{"date":"2026-07-22","tag":"method","text":"Marked the open-weight quality history as a Fable-anchored source series that is rebased in the browser against the selected reference; fixed task-speed remains explicitly Fable-indexed."},{"date":"2026-07-22","tag":"model","text":"Added Gemini 3.6 Flash and Gemini 3.5 Flash-Lite benchmark, pricing, context and throughput evidence from Google and Artificial Analysis; retained missing task-time fields as null."},{"date":"2026-07-21","tag":"model","text":"Added official Kimi K3 API pricing and derived the documented 30/70 workload blend and cost-efficiency score; no task-speed compound is shown because comparable speed evidence remains unavailable."},{"date":"2026-07-19","text":"Added a dated GPU-hosting price tracker with billing-basis normalization, historical Runpod price changes, current Contabo configurations, and quote-aware Infomaniak options. Replaced stale Vast.ai constants with live-marketplace status."},{"date":"2026-07-18","tag":"model","text":"Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5."},{"date":"2026-07-18","tag":"data","text":"Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters."},{"date":"2026-07-18","tag":"data","text":"Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, context, benchmark records and reference eligibility."},{"date":"2026-07-18","tag":"method","text":"Introduced documented quality, speed, cost, SCQ and six-axis capability composites with explicit missing-data rules."},{"date":"2026-07-18","tag":"hardware","text":"Replaced blank hardware fit speeds with supported/unsupported states and evidence-labelled planning estimates."},{"date":"2026-07-18","tag":"routing","text":"Updated recommended routing to Fable for hardest retained-data work, Terra for default engineering and Luna for high-volume subagents."}],"sources":[{"id":"openai-gpt-5-6","title":"GPT-5.6 launch and evaluations","url":"https://openai.com/index/gpt-5-6/","accessed":"2026-07-18"},{"id":"moonshot-kimi-k3-pricing","title":"Kimi K3 API pricing","url":"https://www.kimi.com/resources/kimi-k3-pricing","accessed":"2026-07-21"},{"id":"openai-models","title":"OpenAI model catalog","url":"https://developers.openai.com/api/docs/models","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-terra","title":"GPT-5.6 Terra model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-terra","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-luna","title":"GPT-5.6 Luna model and pricing","url":"https://developers.openai.com/api/docs/models/gpt-5.6-luna","accessed":"2026-08-01"},{"id":"openai-gpt-5-6-price-performance","title":"GPT-5.6 price-performance update","url":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/","accessed":"2026-08-07"},{"id":"openai-gpt-5-6-chat-access","title":"GPT-5.6 Sol improvement and Luna access expansion","url":"https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/","accessed":"2026-08-07"},{"id":"deepseek-v4-flash-0731","title":"DeepSeek V4-Flash-0731 public-beta update","url":"https://api-docs.deepseek.com/updates/","accessed":"2026-08-01"},{"id":"deepseek-v4-pricing","title":"DeepSeek V4 models and pricing","url":"https://api-docs.deepseek.com/quick_start/pricing/","accessed":"2026-08-01"},{"id":"anthropic-models","title":"Claude models overview","url":"https://platform.claude.com/docs/en/about-claude/models/overview","accessed":"2026-07-18"},{"id":"anthropic-retention","title":"Claude API and data retention","url":"https://platform.claude.com/docs/en/manage-claude/api-and-data-retention","accessed":"2026-07-18"},{"id":"google-gemini-3-5","title":"Gemini 3.5 Flash","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash","accessed":"2026-07-18"},{"id":"google-gemini-3-6","title":"Gemini 3.6 Flash model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash","accessed":"2026-07-22"},{"id":"google-gemini-3-5-flash-lite","title":"Gemini 3.5 Flash-Lite model documentation","url":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite","accessed":"2026-07-22"},{"id":"google-flash-july-2026","title":"Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/","accessed":"2026-07-22"},{"id":"aa-gemini-3-6-flash","title":"Artificial Analysis — Gemini 3.6 Flash","url":"https://artificialanalysis.ai/models/gemini-3-6-flash","accessed":"2026-07-22"},{"id":"aa-gemini-3-5-flash-lite","title":"Artificial Analysis — Gemini 3.5 Flash-Lite","url":"https://artificialanalysis.ai/models/gemini-3-5-flash-lite","accessed":"2026-07-22"},{"id":"moonshot-kimi-k3","title":"Kimi K3 technical launch blog","url":"https://www.kimi.com/blog/kimi-k3","accessed":"2026-07-18"},{"id":"moonshot-models","title":"Kimi API model catalog","url":"https://platform.kimi.ai/docs/models","accessed":"2026-07-18"},{"id":"openai-gpt-5-5","title":"Introducing GPT-5.5","url":"https://openai.com/index/introducing-gpt-5-5/","accessed":"2026-07-18"}],"hosting_prices":{"as_of":"2026-07-19","hours_per_month":730,"methodology":"Preserve provider currency and billing basis. normalized_hourly is a comparison-only derivation for fixed monthly plans; monthly_equivalent multiplies hourly rates by 730. It excludes tax, storage, egress, public IPs, support, discounts, utilization, and model throughput. Append observations instead of replacing them.","trend_policy":"Track the same provider, GPU, service tier, region basis, and currency. A changed SKU starts a new series. Two or more observations produce a trend; one observation is a dated baseline.","offers":[{"id":"runpod-h100-sxm-secure","provider":"Runpod","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$2.99/hr","normalized_hourly":2.99,"monthly_equivalent":2182.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":3.99,"source_n":64},{"date":"2026-07-19","value":2.99,"source_n":63}]},{"id":"runpod-a100-sxm-secure","provider":"Runpod","configuration":"1 x A100 SXM","gpu":"A100 SXM","gpu_count":1,"vram_gb":80,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$1.49/hr","normalized_hourly":1.49,"monthly_equivalent":1087.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.94,"source_n":64},{"date":"2026-07-19","value":1.49,"source_n":63}]},{"id":"runpod-l40s-secure","provider":"Runpod","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Multi-region","billing":"Pod, per-second usage","price_basis":"Secure Cloud published rate","currency":"USD","current_price_label":"$0.99/hr","normalized_hourly":0.99,"monthly_equivalent":722.7,"price_status":"published","checked":"2026-07-19","source_n":63,"history":[{"date":"2024-07-12","value":1.19,"source_n":64},{"date":"2026-07-19","value":0.99,"source_n":63}]},{"id":"hyperstack-h100-sxm","provider":"Hyperstack","configuration":"1 x H100 SXM","gpu":"H100 SXM","gpu_count":1,"vram_gb":80,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$2.40/hr","normalized_hourly":2.4,"monthly_equivalent":1752,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":2.4,"source_n":68}]},{"id":"hyperstack-h200-sxm","provider":"Hyperstack","configuration":"1 x H200 SXM","gpu":"H200 SXM","gpu_count":1,"vram_gb":141,"region":"Europe / North America","billing":"On demand, per minute","price_basis":"Published on-demand rate","currency":"USD","current_price_label":"$3.50/hr","normalized_hourly":3.5,"monthly_equivalent":2555,"price_status":"published","checked":"2026-07-19","source_n":68,"history":[{"date":"2026-07-19","value":3.5,"source_n":68}]},{"id":"contabo-l40s-monthly","provider":"Contabo","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€751/mo","normalized_hourly":1.0288,"monthly_equivalent":751,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":1.0288,"source_n":65}]},{"id":"contabo-h100-monthly","provider":"Contabo","configuration":"1 x H100","gpu":"H100","gpu_count":1,"vram_gb":80,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€1,838/mo","normalized_hourly":2.5178,"monthly_equivalent":1838,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.5178,"source_n":65}]},{"id":"contabo-h200-nvl-monthly","provider":"Contabo","configuration":"1 x H200 NVL","gpu":"H200 NVL","gpu_count":1,"vram_gb":141,"region":"Provider location selection","billing":"Fixed monthly plan","price_basis":"Published monthly price, excl. VAT","currency":"EUR","current_price_label":"€2,149/mo","normalized_hourly":2.9438,"monthly_equivalent":2149,"price_status":"published","checked":"2026-07-19","source_n":65,"history":[{"date":"2026-07-19","value":2.9438,"source_n":65}]},{"id":"infomaniak-l40s","provider":"Infomaniak","configuration":"1 x L40S","gpu":"L40S","gpu_count":1,"vram_gb":48,"region":"Switzerland","billing":"Usage-based Public Cloud","price_basis":"Live calculator / availability validation","currency":"CHF","current_price_label":"Calculator / validate","normalized_hourly":null,"monthly_equivalent":null,"price_status":"not_publicly_exposed","checked":"2026-07-19","source_n":66,"history":[]},{"id":"vast-a6000-market","provider":"Vast.ai","configuration":"1 x RTX A6000","gpu":"RTX A6000","gpu_count":1,"vram_gb":48,"region":"Marketplace","billing":"Per second; on-demand, reserved, or interruptible","price_basis":"Live host marketplace","currency":"USD","current_price_label":"Live marketplace","normalized_hourly":null,"monthly_equivalent":null,"price_status":"dynamic_marketplace","checked":"2026-07-19","source_n":69,"history":[]}]}},"model_roster":{"schema_version":"2.0-preview","researched_at":"2026-08-07","default_reference":"fable-5","inclusion_policy":{"default_limit_per_provider":3,"default_scope":"core","summary":"Show a flagship, a balanced or fast model, and one distinctive specialist per provider. Keep siblings and emerging models in expandable extended and watchlist scopes.","promotion_rule":"Promote a watchlist family when at least two signals hold: current official release, meaningful adoption or discussion, differentiated capability or efficiency, and reproducible availability."},"speed_methodology":{"dimensions":["time_to_first_token","output_tokens_per_second","end_to_end_task_time"],"rule":"Compare speed only at the exact model, configuration, provider, region, date, workload, and cache condition. Vendor claims are labelled and unknown values remain unknown.","bands":{"fast":"Designed or measured for low interaction latency","balanced":"General-purpose latency and quality trade-off","deliberate":"Higher reasoning depth or multi-agent execution","unknown":"No comparable evidence"}},"models":[{"id":"fable-5","name":"Claude Fable 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","reference":true,"context":"1M","price":"$10 / $50","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"deliberate","speed_note":"Anthropic comparative latency: slower. No Fable fast mode.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"30-day retention; no zero-data-retention option","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"opus-5","name":"Claude Opus 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Anthropic comparative latency: moderate. Research-preview fast mode claims up to 2.5x higher output throughput at $10 / $50 per MTok.","availability":"GA; API, Bedrock, Google Cloud, Microsoft Foundry","caveat":"Adaptive thinking is on by default; disabling it is supported only through high effort. Independent benchmark results vary materially by effort level and should not be treated as a single generic score.","source":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5"},{"id":"opus-4.8","name":"Claude Opus 4.8","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$5 / $25","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"balanced","speed_note":"Optional API fast mode claims up to 2.5x output throughput at premium pricing.","availability":"GA","source":"https://platform.claude.com/docs/en/build-with-claude/fast-mode"},{"id":"sonnet-5","name":"Claude Sonnet 5","family":"Claude 5","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$3 / $15","license":"Proprietary","control":"effort","levels":["low","medium","high","xhigh","max"],"default_level":"high","speed_band":"fast","speed_note":"Anthropic comparative latency: fast.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"haiku-4.5","name":"Claude Haiku 4.5","family":"Claude 4","provider":"Anthropic","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"200K","price":"$1 / $5","license":"Proprietary","control":"fixed","levels":[],"default_level":"n/a","speed_band":"fast","speed_note":"Anthropic’s latest verified Haiku and fastest listed Claude; no Haiku 5 is in the current official catalog.","availability":"GA","source":"https://platform.claude.com/docs/en/about-claude/models/overview"},{"id":"gpt-5.6-sol","name":"GPT-5.6 Sol","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports max reasoning completes comparable AAII work in 61% less time than Fable 5. API Fast mode, introduced July 30, claims up to 2.5x Standard speed at 2x price with unchanged intelligence; no comparable output-tokens/s figure is published.","availability":"GA; ChatGPT, Codex, API","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-terra","name":"GPT-5.6 Terra","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$2 / $12; cached input $0.20","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 20% from $2.50/$15; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; subscription prices and quota budgets unchanged","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.6-luna","name":"GPT-5.6 Luna","family":"GPT-5.6","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1.05M","price":"$0.20 / $1.20; cached input $0.02","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh","max"],"default_level":"medium","speed_band":"fast","speed_note":"OpenAI positions Luna as the fastest tier and reports coding-agent work in roughly one-third of Fable 5's time; no comparable output-tokens/s figure is published. July 30 list-price cut: 80% from $1/$6; paid Codex and ChatGPT Work usage also consumes fewer credits.","availability":"GA; ChatGPT, Codex, API; rolling out as Free and Go default with unlimited text chats and a Think option","source":"https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/"},{"id":"gpt-5.5","name":"GPT-5.5","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M API / 400K Codex","price":"$5 / $30","license":"Proprietary","control":"reasoning effort","levels":["none","low","medium","high","xhigh"],"default_level":"medium","speed_band":"balanced","speed_note":"OpenAI reports GPT-5.4-class per-token latency; Codex fast mode is 1.5x throughput at 2.5x cost.","availability":"API, ChatGPT and Codex","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-pro","name":"GPT-5.5 Pro","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"commercial_api","scope":"extended","status":"stable","context":"1M","price":"$30 / $180","license":"Proprietary","control":"reasoning effort","levels":["medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Higher-accuracy tier for difficult work; no comparable public output-throughput figure.","availability":"API and eligible ChatGPT plans","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gpt-5.5-instant","name":"GPT-5.5 Instant","family":"GPT-5.5","provider":"OpenAI","region":"US","channel":"harness_bundled","scope":"extended","status":"stable","context":"400K","price":"Included by plan / routed service","license":"Proprietary","control":"fixed","levels":[],"default_level":"provider default","speed_band":"fast","speed_note":"Low-latency ChatGPT route; do not equate service behavior with the reasoning model API.","availability":"ChatGPT","source":"https://openai.com/index/introducing-gpt-5-5/"},{"id":"gemini-3.1-pro","name":"Gemini 3.1 Pro","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"preview","context":"1M","price":"$2 / $12","license":"Proprietary","control":"thinking level","levels":["low","medium","high"],"default_level":"high","speed_band":"deliberate","speed_note":"Google warns high thinking may significantly delay the first answer token.","availability":"Preview","source":"https://ai.google.dev/gemini-api/docs/thinking"},{"id":"gemini-3.6-flash","name":"Gemini 3.6 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$1.50 / $7.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Artificial Analysis measured about 304 output tok/s at high thinking; Google reports fewer reasoning turns and tool calls than 3.5 Flash.","availability":"GA via Gemini API, AI Studio, Gemini app and Antigravity","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"id":"gemini-3.5-flash","name":"Gemini 3.5 Flash","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$1.50 / $9","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"medium","speed_band":"fast","speed_note":"Still offered; migrate new general Flash workloads to 3.6 Flash for lower output price and stronger agentic performance.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash"},{"id":"gemini-3.5-flash-lite","name":"Gemini 3.5 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"1M","price":"$0.30 / $2.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Artificial Analysis measured about 490 output tok/s; optimized for high-volume subagents, document parsing and extraction.","availability":"GA via Gemini API, AI Studio and Gemini app rollout","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"id":"gemini-3.1-flash-lite","name":"Gemini 3.1 Flash-Lite","family":"Gemini 3","provider":"Google","region":"US","channel":"commercial_api","scope":"extended","status":"legacy","context":"1M","price":"$0.25 / $1.50","license":"Proprietary","control":"thinking level","levels":["minimal","low","medium","high"],"default_level":"minimal","speed_band":"fast","speed_note":"Still offered; Gemini 3.5 Flash-Lite is the current high-throughput migration target.","availability":"Stable legacy option","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite"},{"id":"grok-4.5","name":"Grok 4.5","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"core","status":"stable","context":"500K","price":"$2 / $6","license":"Proprietary","control":"reasoning effort","levels":["low","medium","high"],"default_level":"high","speed_band":"fast","speed_note":"xAI reports 80 output tokens/s; vendor measurement conditions are not fully comparable here.","availability":"API and integrations; verify regional availability","source":"https://x.ai/news/grok-4-5"},{"id":"grok-4.20-multi-agent","name":"Grok 4.20 Multi-Agent","family":"Grok 4","provider":"xAI","region":"US","channel":"commercial_api","scope":"extended","status":"specialist","context":"1M","price":"$1.25 / $2.50","license":"Proprietary","control":"agent count","levels":["low","medium","high","xhigh"],"default_level":"high","speed_band":"deliberate","speed_note":"Levels control a 4- or 16-agent process, not ordinary reasoning depth.","availability":"API","source":"https://docs.x.ai/developers/model-capabilities/text/multi-agent"},{"id":"composer-2.5","name":"Composer 2.5","family":"Composer","provider":"Cursor","region":"US","channel":"harness_bundled","scope":"core","status":"stable","context":"Managed by Cursor","price":"$0.50 / $2.50 standard","license":"Proprietary","control":"speed variant","levels":["standard","fast"],"default_level":"fast","speed_band":"fast","speed_note":"Fast is the default; Cursor documents no public low/medium/high effort selector.","availability":"Cursor only","source":"https://cursor.com/blog/composer-2-5"},{"id":"kimi-k3","name":"Kimi K3","family":"Kimi K3","provider":"Moonshot AI","region":"China","channel":"commercial_api_open_weights_announced","scope":"core","status":"stable","context":"1M","price":"$3 / $15; cached input $0.30","license":"Terms pending weight release","control":"reasoning effort","levels":["max"],"default_level":"max","speed_band":"deliberate","speed_note":"No comparable K3 output-rate figure at launch; max thinking is always enabled.","availability":"API, Kimi, Kimi Work and Kimi Code; weights promised by July 27","source":"https://www.kimi.com/resources/kimi-k3-pricing"},{"id":"kimi-k2.7-code","name":"Kimi K2.7 Code","family":"Kimi K2.7","provider":"Moonshot AI","region":"China","channel":"commercial_api","scope":"core","status":"specialist","context":"256K","price":"Current API pricing","license":"Proprietary API","control":"service variant","levels":["standard","high-speed"],"default_level":"standard","speed_band":"fast","speed_note":"High-speed service is documented at about 180 tok/s and up to 260 tok/s for short contexts.","availability":"Kimi API and Kimi Code","source":"https://platform.kimi.ai/docs/models"},{"id":"deepseek-v4","name":"DeepSeek V4","family":"DeepSeek V4","provider":"DeepSeek","region":"China","channel":"open_weight","scope":"core","status":"public_beta","context":"1M","price":"Flash $0.14 / $0.28; cached input $0.0028","license":"MIT","control":"mode","levels":["non-think","think","max"],"default_level":"think","speed_band":"balanced","speed_note":"V4-Flash-0731 reports 82.7 on Terminal-Bench 2.1 at max effort in DeepSeek Harness minimal mode; comparable latency remains unknown.","availability":"V4-Flash-0731 API public beta and weights; future 2x peak-hours pricing announced without an effective date","variants":["Pro","Flash"],"source":"https://api-docs.deepseek.com/updates/"},{"id":"qwen-3.6","name":"Qwen3.6","family":"Qwen3.6","provider":"Alibaba Qwen","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"mode and token budget","levels":["non-thinking","thinking","budget"],"default_level":"thinking","speed_band":"balanced","speed_note":"Budget is provider-native; do not translate it into invented effort labels.","availability":"API and weights","variants":["27B","35B-A3B"],"source":"https://github.com/QwenLM/Qwen3.6"},{"id":"glm-5.2","name":"GLM-5.2","family":"GLM-5","provider":"Z.ai","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"MIT","control":"effort","levels":["adaptive"],"default_level":"adaptive","speed_band":"unknown","speed_note":"Architecture efficiency claims are not a comparable API latency measurement.","availability":"API and weights","source":"https://z.ai/blog/glm-5.2"},{"id":"kimi-k2.5","name":"Kimi K2.5","family":"Kimi K2","provider":"Moonshot AI","region":"China","channel":"open_weight","scope":"extended","status":"deprecated","context":"256K","price":"API and self-hosted","license":"Modified MIT","control":"mode","levels":["instant","thinking"],"default_level":"thinking","speed_band":"balanced","speed_note":"No comparable official throughput measurement; Instant is the lower-latency mode.","availability":"Existing users only; platform sunset August 31, 2026","source":"https://platform.kimi.ai/docs/models"},{"id":"minimax-m3","name":"MiniMax M3","family":"MiniMax M3","provider":"MiniMax","region":"China","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"API and self-hosted","license":"Community; non-commercial by default","control":"thinking mode","levels":["disabled","adaptive","enabled"],"default_level":"adaptive","speed_band":"balanced","speed_note":"Vendor claims decode improvements versus M2, not cross-provider latency.","availability":"API and weights","caveat":"License is not permissive open source","source":"https://www.minimax.io/blog/minimax-m3"},{"id":"mistral-small-4","name":"Mistral Small 4","family":"Mistral 3/4","provider":"Mistral AI","region":"EU","channel":"open_weight","scope":"core","status":"stable","context":"256K","price":"API and self-hosted","license":"Apache-2.0","control":"reasoning effort","levels":["none","high"],"default_level":"none","speed_band":"fast","speed_note":"Vendor claims 40% lower completion time and 3x requests/s versus Small 3.","availability":"API and weights","source":"https://mistral.ai/it/news/mistral-small-4/"},{"id":"llama-4","name":"Llama 4","family":"Llama 4","provider":"Meta","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"10M advertised for Scout","price":"Self-hosted","license":"Llama 4 Community License","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Performance depends on deployment; no comparable official API latency.","availability":"Weights","variants":["Scout","Maverick"],"caveat":"Not OSI-open; EU multimodal and large-platform restrictions apply","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"id":"gemma-3","name":"Gemma 3","family":"Gemma 3","provider":"Google","region":"US","channel":"open_weight","scope":"extended","status":"stable","context":"128K","price":"Self-hosted","license":"Google Gemma Terms","control":"none documented","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"Deployment dependent; optimized for edge and local use.","availability":"Weights","variants":["1B","4B","12B","27B"],"source":"https://ai.google.dev/gemma/docs/core/model_card_3"},{"id":"gpt-oss","name":"gpt-oss","family":"gpt-oss","provider":"OpenAI","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"128K","price":"Self-hosted","license":"Apache-2.0 plus usage policy","control":"reasoning effort","levels":["low","medium","high"],"default_level":"medium","speed_band":"unknown","speed_note":"Depends on hardware and serving stack.","availability":"Weights","variants":["120b","20b"],"source":"https://openai.com/index/introducing-gpt-oss/"},{"id":"apertus-v1.1-4b-instruct","name":"Apertus v1.1 4B Instruct","family":"Apertus v1.1","provider":"Swiss AI Initiative","region":"Switzerland","channel":"open_weight","scope":"core","status":"stable","context":"4K","price":"Self-hosted","license":"Apache-2.0","control":"sampling parameters","levels":[],"default_level":"n/a","speed_band":"unknown","speed_note":"No comparable throughput measurement is published; performance depends on hardware, runtime and quantization.","availability":"Weights; Transformers, vLLM, SGLang and quantized checkpoints","variants":["0.5B","1.5B","4B"],"source":"https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct"},{"id":"inkling","name":"Inkling","family":"Inkling","provider":"Thinking Machines Lab","region":"US","channel":"open_weight","scope":"core","status":"stable","context":"1M","price":"No metered API; open weights + Tinker fine-tuning platform","license":"Apache 2.0","control":"n/a","levels":[],"default_level":null,"speed_band":"unknown","speed_note":"No published throughput figure; not yet independently benchmarked for this roster.","availability":"Open weights (Hugging Face), Tinker fine-tuning platform, third-party inference providers","source":"https://thinkingmachines.ai/model-card/inkling/"}],"watchlist":[{"name":"Step 3.5 Flash","provider":"StepFun","region":"China","license":"Apache-2.0","signal":"Official 100-300 output tok/s claim","source":"https://github.com/stepfun-ai/Step-3.5-Flash"},{"name":"MiMo-V2-Flash","provider":"Xiaomi","region":"China","license":"Apache-2.0","signal":"Thinking toggle and generation-speed claim","source":"https://github.com/XiaomiMiMo/MiMo-V2-Flash"},{"name":"Seed-OSS-36B","provider":"ByteDance","region":"China","license":"Apache-2.0","signal":"512K context and controllable thinking budget","source":"https://github.com/ByteDance-Seed/seed-oss"},{"name":"Hy3 Preview","provider":"Tencent","region":"China","license":"Tencent Hy Community License","signal":"Preview multimodal MoE family","source":"https://github.com/Tencent-Hunyuan/Hy3-preview"}]}}&lt;/script&gt;
&lt;section class="tab-panel" id="tab-sources"&gt;
 &lt;div id="sources-content"&gt;&lt;/div&gt;
 &lt;/section&gt;
&lt;/div&gt;</description></item><item><title>Corrected the GPT-5.6 pricing history from OpenAI's July 30 announcement: Terra fell...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-08-07-corrected-the-gpt-5-6-pricing-history-from-openai-s-july-30/</link><pubDate>Fri, 07 Aug 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-08-07-corrected-the-gpt-5-6-pricing-history-from-openai-s-july-30/</guid><description>&lt;p&gt;&lt;strong&gt;Pricing&lt;/strong&gt; · Corrected the GPT-5.6 pricing history from OpenAI&amp;rsquo;s July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. The report now records lower paid Codex and ChatGPT Work credit consumption, unchanged subscription prices and quota budgets, Sol API Fast mode (up to 2.5× Standard speed at 2× price), and the August 6 ChatGPT Luna access expansion.&lt;/p&gt;</description></item><item><title>Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-08-01-updated-deepseek-v4-flash-to-the-0731-public-beta-with-1m-co/</link><pubDate>Sat, 01 Aug 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-08-01-updated-deepseek-v4-flash-to-the-0731-public-beta-with-1m-co/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output, thinking/non-thinking modes, $0.14/$0.28 per-million pricing, $0.0028 cached input, and vendor-reported agent results including Terminal-Bench 2.1 82.7 and Agents&amp;rsquo; Last Exam 25.2.&lt;/p&gt;</description></item><item><title>Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-08-01-updated-gpt-5-6-terra-to-2-12-per-million-input-output-token/</link><pubDate>Sat, 01 Aug 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-08-01-updated-gpt-5-6-terra-to-2-12-per-million-input-output-token/</guid><description>&lt;p&gt;&lt;strong&gt;Pricing&lt;/strong&gt; · Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to $0.20/$1.20 from OpenAI&amp;rsquo;s live model pages. Recomputed the tracked 30/70 workload blends and effort-burn ratios. Superseded on August 7 with OpenAI&amp;rsquo;s published July 30 effective date and full price-change details.&lt;/p&gt;</description></item><item><title>Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-28-added-thinking-machines-lab-and-its-first-model-inkling-a-97/</link><pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-28-added-thinking-machines-lab-and-its-first-model-inkling-a-97/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B active) open-weight (Apache 2.0) multimodal MoE with 1M-token context, released 2026-07-15. No first-party per-token API pricing is published (monetized via the Tinker fine-tuning platform); benchmark scores are published on the model card but not yet normalized into this roster&amp;rsquo;s comparable set, so quality/speed/cost fields are left unknown rather than estimated.&lt;/p&gt;</description></item><item><title>Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-25-added-claude-opus-5-general-availability-1m-context-128k-max/</link><pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-25-added-claude-opus-5-general-availability-1m-context-128k-max/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive thinking by default, $5/$25 per MTok base pricing, and official cloud-platform availability. Added independent Artificial Analysis evidence (61 Intelligence Index at max effort; 52.3 output tok/s) with effort-specific caveats. Gemini 3.5 Flash Cyber remains limited to CodeMender government and trusted-partner pilots; GPT-Live and Muse Spark 1.1 remain non-API products, so none were added to the API roster. Claude Opus 4.7 Fast Mode was removed July 24.&lt;/p&gt;</description></item><item><title>Refreshed current-source evidence: added OpenAI's GPT-5.6 launch table, Google's...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-24-refreshed-current-source-evidence-added-openai-s-gpt-5-6-lau/</link><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-24-refreshed-current-source-evidence-added-openai-s-gpt-5-6-lau/</guid><description>&lt;p&gt;&lt;strong&gt;Benchmarks&lt;/strong&gt; · Refreshed current-source evidence: added OpenAI&amp;rsquo;s GPT-5.6 launch table, Google&amp;rsquo;s Managed Agents update, and Scale&amp;rsquo;s public SWE-bench Pro leaderboard. Clarified that vendor launch tables and the public leaderboard are not directly comparable because their model versions and harnesses differ.&lt;/p&gt;</description></item><item><title>Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-22-added-google-s-ga-gemini-3-6-flash-and-gemini-3-5-flash-lite/</link><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-22-added-google-s-ga-gemini-3-6-flash-and-gemini-3-5-flash-lite/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M context, 64K output, current API pricing, Artificial Analysis intelligence and throughput measurements, and Google’s published coding and agentic benchmarks.&lt;/p&gt;</description></item><item><title>Made the selected reference propagate through open-weight quality headlines...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-22-made-the-selected-reference-propagate-through-open-weight-qu/</link><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-22-made-the-selected-reference-propagate-through-open-weight-qu/</guid><description>&lt;p&gt;&lt;strong&gt;Fix&lt;/strong&gt; · Made the selected reference propagate through open-weight quality headlines, market-signal history, comparison headings, model analytics, self-hosting quality and capability market position. Replaced the model and harness inventory mini-lines with stacked proprietary/open category areas based on release-tag roster snapshots.&lt;/p&gt;</description></item><item><title>Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-22-recorded-google-s-broader-flash-shift-gemini-3-5-flash-cyber/</link><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-22-recorded-google-s-broader-flash-shift-gemini-3-5-flash-cyber/</guid><description>&lt;p&gt;&lt;strong&gt;Data&lt;/strong&gt; · Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to governments and trusted CodeMender partners; Gemini Omni Flash and Nano Banana 2 Lite remain specialized media models rather than general-purpose roster entries.&lt;/p&gt;</description></item><item><title>Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-21-added-apertus-v1-1-4b-instruct-the-largest-newly-released-ap/</link><pubDate>Tue, 21 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-21-added-apertus-v1-1-4b-instruct-the-largest-newly-released-ap/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint: fully open Apache 2.0 weights and data, 4K context, 1.7T-token distillation, 1,811 languages, and official BF16, FP8, NVFP4A16, INT3, INT4 and INT6 variants.&lt;/p&gt;</description></item><item><title>Added Moonshot's official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-21-added-moonshot-s-official-kimi-k3-api-pricing-3-m-uncached-i/</link><pubDate>Tue, 21 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-21-added-moonshot-s-official-kimi-k3-api-pricing-3-m-uncached-i/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Added Moonshot&amp;rsquo;s official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached input, and $15/M output; the 30/70 workload blend is $11.40/M before reasoning-effort effects.&lt;/p&gt;</description></item><item><title>Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-21-refreshed-five-open-agent-harnesses-from-their-canonical-git/</link><pubDate>Tue, 21 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-21-refreshed-five-open-agent-harnesses-from-their-canonical-git/</guid><description>&lt;p&gt;&lt;strong&gt;Agent harnesses&lt;/strong&gt; · Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI 0.144.6, Gemini CLI 0.51.0, OpenCode 1.18.4, Cline 4.0.10, and Goose 1.43.0; updated repository star snapshots and notable release capabilities.&lt;/p&gt;</description></item><item><title>Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-18-added-kimi-k3-from-moonshot-primary-sources-2-8t-sparse-moe/</link><pubDate>Sat, 18 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-18-added-kimi-k3-from-moonshot-primary-sources-2-8t-sparse-moe/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5.&lt;/p&gt;</description></item><item><title>Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-18-added-source-grounded-gpt-5-6-sol-terra-and-luna-pricing-1-0/</link><pubDate>Sat, 18 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-18-added-source-grounded-gpt-5-6-sol-terra-and-luna-pricing-1-0/</guid><description>&lt;p&gt;&lt;strong&gt;Data&lt;/strong&gt; · Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark registers and selectable reference configurations. Added documented quality, speed, cost and capability composites in data/report-metrics.json.&lt;/p&gt;</description></item><item><title>Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-18-restored-gpt-5-5-family-visibility-verified-haiku-4-5-as-the/</link><pubDate>Sat, 18 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-18-restored-gpt-5-5-family-visibility-verified-haiku-4-5-as-the/</guid><description>&lt;p&gt;&lt;strong&gt;Data&lt;/strong&gt; · Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters.&lt;/p&gt;</description></item><item><title>Updated the action queue and recommended routing: Fable for hardest retained-data...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-18-updated-the-action-queue-and-recommended-routing-fable-for-h/</link><pubDate>Sat, 18 Jul 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-07-18-updated-the-action-queue-and-recommended-routing-fable-for-h/</guid><description>&lt;p&gt;&lt;strong&gt;Routing&lt;/strong&gt; · Updated the action queue and recommended routing: Fable for hardest retained-data workloads, Terra for default engineering, Luna for high-volume subagents, with explicit escalation rules.&lt;/p&gt;</description></item><item><title>Dashboard market sweep v1.5.0: real-world re-grounding</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-06-dashboard-market-sweep-v1-5-0-real-world-re-grounding/</link><pubDate>Sat, 06 Jun 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-06-dashboard-market-sweep-v1-5-0-real-world-re-grounding/</guid><description>&lt;p&gt;&lt;strong&gt;Policy&lt;/strong&gt; · Dashboard market sweep v1.5.0: real-world re-grounding. Replaced fictional Mythos/GPT-5.5-Cyber rows with verified models; added Nvidia Nemotron coalition, Kimi K2.6, GLM-5, Cohere Command A+, SubQ 1M-Preview.&lt;/p&gt;</description></item><item><title>Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-06-restored-gpt-5-3-codex-spark-feb-12-2026-release-chatgpt-pro/</link><pubDate>Sat, 06 Jun 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-06-restored-gpt-5-3-codex-spark-feb-12-2026-release-chatgpt-pro/</guid><description>&lt;p&gt;&lt;strong&gt;Fix&lt;/strong&gt; · Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K context, 1000+ tok/s on Cerebras) and Hermes Agent v0.16.0 (Nous Research, MIT, self-hosted multi-platform agent) — both were incorrectly removed in v1.5.0 sweep.&lt;/p&gt;</description></item><item><title>Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-04-nvidia-nemotron-coalition-formed-black-forest-labs-cursor-la/</link><pubDate>Thu, 04 Jun 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-04-nvidia-nemotron-coalition-formed-black-forest-labs-cursor-la/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral, Perplexity, Reflection AI, Sarvam, Thinking Machines Lab as inaugural members.&lt;/p&gt;</description></item><item><title>Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context...</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-04-nvidia-releases-nemotron-3-ultra-550b-55b-moe-hybrid-mamba-t/</link><pubDate>Thu, 04 Jun 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-04-nvidia-releases-nemotron-3-ultra-550b-55b-moe-hybrid-mamba-t/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context, NVIDIA Open Model License) at Computex — first frontier-scale open model from Nvidia.&lt;/p&gt;</description></item><item><title>Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model</title><link>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-01-nvidia-cosmos-3-launched-open-physical-ai-robotics-foundatio/</link><pubDate>Mon, 01 Jun 2026 00:00:00 +0000</pubDate><guid>https://projectious-work.github.io/ai-market-research/blog/2026/2026-06-01-nvidia-cosmos-3-launched-open-physical-ai-robotics-foundatio/</guid><description>&lt;p&gt;&lt;strong&gt;Model roster&lt;/strong&gt; · Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model.&lt;/p&gt;</description></item></channel></rss>