Worth reading
GPT-6 Astra: staged access and a 272K price cliff
GPT-6 Astra brings a 1.05M-token context window and a new long-context price tier. Check access and the 272K threshold before budgeting a migration.
Sourced capabilities, source-linked model records, and your own tests — in one place.
One clear model decision
Three steps. No universal winner.
Worth reading
GPT-6 Astra brings a 1.05M-token context window and a new long-context price tier. Check access and the 272K threshold before budgeting a migration.
New and changed
GPT-6 Astra brings a 1.05M-token context window and a new long-context price tier. Check access and the 272K threshold before budgeting a migration.
DeepSeek raised V4 API prices on August 16, 2026 and split every rate into peak and off-peak halves. V4-Flash output went from $0.28 to $1.32 during peak hours.
Gemini 3.7 Flash went GA on August 13, 2026 at $0.75/$3.75 per 1M tokens, the same rate 3.6 Flash was cut to. Both return to $1.50/$7.50 on January 1, 2027.
Z.AI's API release adds post-training gains, always-on reasoning, and three compatible protocols at $1.40/$4.40.
Google's stable multimodal model starts at $0.75/$3.75 through 2026, with a 1M-token window and a dated price change.
xAI targets long coding runs and interactive agents at $2/$6, while its docs publish no numeric text-output cap.
Grok 4.5 review: xAI's coding specialist, built with Cursor, priced at $2/$6, with a smaller 500K context window than Grok 4.3 despite what some blogs claim.
GLM-5.2 review: Zhipu's MIT-licensed, open-weight 753B model at $1.40/$4.40 per million tokens, with coding benchmarks Z.AI says beat GPT-5.5 and Opus 4.7.
Anthropic's current Opus at the same $5/$25 rate — what changes over 4.8, and who should move.
A candidate for visual and structured-output work based on OpenAI's published positioning; validate it on your own prompts.
A retired multimodal model that Google deprecated in March 2026 in favor of Gemini 3.1 Pro.
Cursor, GitHub Copilot, Windsurf, and Cody on the same Markdown-exporter task.
ElevenLabs, OpenAI Whisper, and Cartesia Sonic on latency, accuracy, and naturalness.
Advertised context capacity versus reproducible retrieval checks you can run on your own documents.
The cheapest model for chat, coding, RAG, agents, classification, and summarization.
Where Llama 4, Mistral Large 3, DeepSeek-V4, and Qwen 3.6 stand against the closed labs.
MMLU is saturated above 90%. The benchmarks worth tracking now.
Claude Opus 4.7, GPT-5, and Gemini 3.1 Pro Preview: sourced specifications, editorial tradeoffs, and a local evaluation plan.
Llama 4, Mistral, DeepSeek, Qwen, the small-model tier, and what it takes to self-host.
What AI costs by model and workload, and where teams overspend.
An interactive comparison covering source-linked pricing, named benchmark records, context windows, and documented capabilities for all 36 models in the shared index.
All 36 factual model records, sortable by published fields →
All 114 articles in the archive →
Provider guides: OpenAI, Anthropic, Google, and open weights →
benchr is an evidence-led reference, not a generic listicle. Figures are identified as official provider data, third-party benchmark results, or benchr editorial estimates so readers can judge the evidence behind each claim.