This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

News

Dated changelog of Signal Room data updates: model roster changes, benchmark refreshes, and fixes.

Every entry here is generated from the same changelog that drives the Signal Room report itself (data/market-state.json) – model roster additions, benchmark refreshes, and fixes, in the order they happened.

Corrected the GPT-5.6 pricing history from OpenAI's July 30 announcement: Terra fell...

Pricing · Corrected the GPT-5.6 pricing history from OpenAI’s July 30 announcement: Terra fell 20% from $2.50/$15 to $2/$12 per MTok and Luna fell 80% from $1/$6 to $0.20/$1.20. The report now records lower paid Codex and ChatGPT Work credit consumption, unchanged subscription prices and quota budgets, Sol API Fast mode (up to 2.5× Standard speed at 2× price), and the August 6 ChatGPT Luna access expansion.

Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output...

Model roster · Updated DeepSeek V4-Flash to the 0731 public beta with 1M context, 384K max output, thinking/non-thinking modes, $0.14/$0.28 per-million pricing, $0.0028 cached input, and vendor-reported agent results including Terminal-Bench 2.1 82.7 and Agents’ Last Exam 25.2.

Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to...

Pricing · Updated GPT-5.6 Terra to $2/$12 per million input/output tokens and Luna to $0.20/$1.20 from OpenAI’s live model pages. Recomputed the tracked 30/70 workload blends and effort-burn ratios. Superseded on August 7 with OpenAI’s published July 30 effective date and full price-change details.

Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B...

Model roster · Added Thinking Machines Lab and its first model, Inkling: a 975B-parameter (41B active) open-weight (Apache 2.0) multimodal MoE with 1M-token context, released 2026-07-15. No first-party per-token API pricing is published (monetized via the Tinker fine-tuning platform); benchmark scores are published on the model card but not yet normalized into this roster’s comparable set, so quality/speed/cost fields are left unknown rather than estimated.

Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive...

Model roster · Added Claude Opus 5: general availability, 1M context, 128K maximum output, adaptive thinking by default, $5/$25 per MTok base pricing, and official cloud-platform availability. Added independent Artificial Analysis evidence (61 Intelligence Index at max effort; 52.3 output tok/s) with effort-specific caveats. Gemini 3.5 Flash Cyber remains limited to CodeMender government and trusted-partner pilots; GPT-Live and Muse Spark 1.1 remain non-API products, so none were added to the API roster. Claude Opus 4.7 Fast Mode was removed July 24.

Refreshed current-source evidence: added OpenAI's GPT-5.6 launch table, Google's...

Benchmarks · Refreshed current-source evidence: added OpenAI’s GPT-5.6 launch table, Google’s Managed Agents update, and Scale’s public SWE-bench Pro leaderboard. Clarified that vendor launch tables and the public leaderboard are not directly comparable because their model versions and harnesses differ.

Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M...

Model roster · Added Google’s GA Gemini 3.6 Flash and Gemini 3.5 Flash-Lite with stable model IDs, 1M context, 64K output, current API pricing, Artificial Analysis intelligence and throughput measurements, and Google’s published coding and agentic benchmarks.

Made the selected reference propagate through open-weight quality headlines...

Fix · Made the selected reference propagate through open-weight quality headlines, market-signal history, comparison headings, model analytics, self-hosting quality and capability market position. Replaced the model and harness inventory mini-lines with stacked proprietary/open category areas based on release-tag roster snapshots.

Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to...

Data · Recorded Google’s broader Flash shift: Gemini 3.5 Flash Cyber remains restricted to governments and trusted CodeMender partners; Gemini Omni Flash and Nano Banana 2 Lite remain specialized media models rather than general-purpose roster entries.

Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint...

Model roster · Added Apertus-v1.1-4B-Instruct, the largest newly released Apertus Mini checkpoint: fully open Apache 2.0 weights and data, 4K context, 1.7T-token distillation, 1,811 languages, and official BF16, FP8, NVFP4A16, INT3, INT4 and INT6 variants.

Added Moonshot's official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached...

Model roster · Added Moonshot’s official Kimi K3 API pricing: $3/M uncached input, $0.30/M cached input, and $15/M output; the 30/70 workload blend is $11.40/M before reasoning-effort effects.

Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI...

Agent harnesses · Refreshed five open agent harnesses from their canonical GitHub releases: Codex CLI 0.144.6, Gemini CLI 0.51.0, OpenCode 1.18.4, Cline 4.0.10, and Goose 1.43.0; updated repository star snapshots and notable release capabilities.

Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native...

Model roster · Added Kimi K3 from Moonshot primary sources: 2.8T sparse MoE, 1M context, native vision, max-only thinking at launch, API availability, and a vendor-suite quality comparison against Fable 5.

Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark...

Data · Added source-grounded GPT-5.6 Sol, Terra and Luna pricing, 1.05M context, benchmark registers and selectable reference configurations. Added documented quality, speed, cost and capability composites in data/report-metrics.json.

Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku...

Data · Restored GPT-5.5 family visibility, verified Haiku 4.5 as the latest public Haiku, replaced quality compound display with quality vs selected reference, and removed non-actionable headline cost/policy counters.

Updated the action queue and recommended routing: Fable for hardest retained-data...

Routing · Updated the action queue and recommended routing: Fable for hardest retained-data workloads, Terra for default engineering, Luna for high-volume subagents, with explicit escalation rules.

Dashboard market sweep v1.5.0: real-world re-grounding

Policy · Dashboard market sweep v1.5.0: real-world re-grounding. Replaced fictional Mythos/GPT-5.5-Cyber rows with verified models; added Nvidia Nemotron coalition, Kimi K2.6, GLM-5, Cohere Command A+, SubQ 1M-Preview.

Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K...

Fix · Restored GPT-5.3-Codex-Spark (Feb 12, 2026 release; ChatGPT Pro research preview, 128K context, 1000+ tok/s on Cerebras) and Hermes Agent v0.16.0 (Nous Research, MIT, self-hosted multi-platform agent) — both were incorrectly removed in v1.5.0 sweep.

Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral...

Model roster · Nvidia Nemotron Coalition formed: Black Forest Labs, Cursor, LangChain, Mistral, Perplexity, Reflection AI, Sarvam, Thinking Machines Lab as inaugural members.

Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context...

Model roster · Nvidia releases Nemotron 3 Ultra (550B/55B MoE, hybrid Mamba-Transformer, 1M context, NVIDIA Open Model License) at Computex — first frontier-scale open model from Nvidia.

Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model

Model roster · Nvidia Cosmos 3 launched — open physical-AI / robotics foundation model.