This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

Signal Room

The current market report – model roster, configurations, speed evidence, agent harnesses, and self-hosting economics.

Signal Room tracks the current model roster, provider-native reasoning configurations, speed evidence, agent harnesses, and self-hosting economics – with evidence classes and source links that stay visible so you can make your own tradeoffs. The reference model, jurisdiction, and workload selections in the bar above carry across every page below.

For the “00 Now” snapshot, see the site landing page. For the formulas and evidence classes behind every number here, see the Data Methodology.

SectionCovers
01 Market & EconomicsCurrent roster, benchmark register, capability radar, quota-burn cross-matrix, subscription tiers, agent policy
02 ToolsAgent harness landscape and detailed profiles
03 InfrastructureHardware x model fit, hardware options, hosting price tracker, inference frameworks
04 DecisionsCurrent recommendation, routing strategy, decision matrix
05 EvidenceEvery cited source, grouped by category

1 - 01 Market & Economics

Current roster, benchmark register, capability radar, quota-burn cross-matrix, subscription tiers, and agent policy.

Current roster

A deliberately compact provider roster: flagship, balanced or fast, and differentiated specialist models. Search, filter, or sort the columns; linked values open the primary provider evidence.

ModelProviderScopeStatusControl / defaultAvailable levelsSpeedSpeed evidenceAPI input / outputContextRegionAvailability

Benchmark register

Quality % is computed against the selected reference (top-right). SWE-bench Pro is the trustworthy benchmark; Verified is contaminated.

ModelProviderQuality vs refSWE-ProSWE-VerLCBAIMEIn $/MtokOut $/MtokCacheContexttok/sReleased

Capability radar

Compare six editorial capability axes for two selectable models. Choose a focus to increase that axis's weight in the capability score; the quality column remains the separate benchmark-derived value relative to the selected reference.

Quota burn cross-matrix

Burn ratios shown as multiples of the selected reference model at medium effort. OpenAI (Codex /effort) and Anthropic (Claude Code /effort) share the same vocabulary — low / medium / high / xhigh — with Anthropic adding 'max' for Opus 4.7. Google uses thinking budgets. Multipliers stack: Fast mode ×2.0, cached input ×0.6, plan mode forces high. Switch the reference dropdown at the top of the page to recompute all ratios.

Subscription tiers

ProviderTier$/moLimitsModelsFeatures

Automated agent policy

ProviderSub on automationEnforcementFirst-party exceptionAPI needed?

2 - 02 Tools

Agent harness landscape and detailed profiles.

Agent harness landscape

Two workflow modes: supervised (you wait) and autonomous (overnight). Anthropic OAuth blocked third-party tools on April 4, 2026 — affected harnesses marked with *.

HarnessVendorCategoryMCPSkillsHooksSubagentsVoiceRemoteComputer UseLSPMemorySWE-ProPricing

Detailed profiles

3 - 03 Infrastructure

Hardware x model fit, hardware options, hosting price tracker, and inference frameworks.

Hardware × model fit

Quantization, VRAM used, and tok/s estimates per hardware config. Quality % is vs selected reference.

Hardware options

HardwareTypeVRAMCostNotes

Hosting price tracker

Published prices by provider-native billing basis. The 730-hour normalization is a comparison aid, not an effective token cost or a quoted monthly bill.

Provider / configurationVRAMBilling basisPublished priceNormalized hourly730h equivalentPrice history

Inference frameworks

FrameworkBest forNotes

4 - 04 Decisions

Current recommendation, routing strategy, and decision matrix.

5 - 05 Evidence

Every cited source, grouped by category.