The decision platform for AI models

Source-backed model intelligence, your own evaluation cases, production cost scenarios, and migration planning — connected in one workspace.

AI model intelligence workspace

Explore. Test. Decide. Ship.

Go from scattered model claims to a decision you can defend. Search verified model and API data, build a shortlist, run your own evaluation set, estimate the bill, and plan migrations before a model retires.

27verified models
25logged changes
41retirement records
4 labelsevidence categories

Live record

Latest model changes

View all changes →
LoadingReading the verified change record…

Latest

  1. 01

    GPT-5.6: Sol, Terra, Luna

    Sol, Terra, and Luna reached general availability on July 9. Sol now lists a 1.05M context window and $5/$30 standard pricing; the article preserves the preview-to-GA correction trail.

  2. 02

    Claude Fable 5

    The $10/$50 Mythos-class model is available worldwide again after its June 12–30 suspension; access was restored July 1. Safety-classified requests refuse explicitly, with another-model retry only when an app configures it.

  3. 03

    Claude Opus 4.8

    Same price as 4.7, a small leaderboard bump, and an honesty gain that catches its own bugs.

  4. 04

    GPT-5.5

    OpenAI's gains land in agentic coding and computer use, at roughly double GPT-5's API price.

  5. 05

    Opus 4.8 vs GPT-5.5

    Both list $5 per million input tokens. Compare their published specifications, then validate the fit on your own stack.

  6. 06

    Claude Mythos

    Anthropic's restricted frontier model, gated under Project Glasswing. What it is and why it's locked away.

Reviews

  1. 07

    Claude Opus 4.7

    Strong on architecture decisions and long-document analysis. Vision is its weaker side.

  2. 08

    GPT-5

    A candidate for visual and structured-output work based on OpenAI's published positioning; validate it on your own prompts.

  3. 09

    Gemini 3 Pro

    A retired multimodal model that Google deprecated in March 2026 in favor of Gemini 3.1 Pro.

Comparisons

  1. 10

    Coding assistants

    Cursor, GitHub Copilot, Windsurf, and Cody on the same Markdown-exporter task.

  2. 11

    Voice models

    ElevenLabs, OpenAI Whisper, and Cartesia Sonic on latency, accuracy, and naturalness.

  3. 12

    Context windows

    Advertised context capacity versus reproducible retrieval checks you can run on your own documents.

Analysis

  1. 13

    Price per use case

    The cheapest model for chat, coding, RAG, agents, classification, and summarization.

  2. 14

    The open-weight tier

    Where Llama 4, Mistral Large 2, DeepSeek-V3.1, and Qwen 3 stand against the closed labs.

  3. 15

    Why benchmarks stopped telling you

    MMLU is saturated above 90%. The benchmarks worth tracking now.

Guides

  1. 16

    Frontier models

    Claude Opus 4.7, GPT-5, and Gemini 3.1 Pro Preview: sourced specifications, editorial tradeoffs, and a local evaluation plan.

  2. 17

    Open-weight models

    Llama 4, Mistral, DeepSeek, Qwen, the small-model tier, and what it takes to self-host.

  3. 18

    AI costs

    What AI costs by model and workload, and where teams overspend.

Compare models directly

An interactive comparison covering pricing, benchmarks, context windows, and capability ratings for all 27 models in the shared index, across frontier, mid, and open-weight tiers.

Open the comparison tool

All 27 models ranked by benchr Rating →

All 88 articles in the archive →

Recent model releases →

benchr is an evidence-led reference, not a generic listicle. Figures are identified as official provider data, third-party benchmark results, or benchr editorial estimates so readers can judge the evidence behind each claim.

Updates

Follow new pieces through RSS, recent releases, and the changelog.