skip to content
#ai
aihackers.net practical notes on building with AI
Latest Models Tools AI Value Build Safely
Search

Benchmarks

tag: Benchmarks

  • 2026-09-27 | GPT-6.1 Sol, Sol and Luna: Pricing and Evidence Compare GPT-6.1 Sol API rates and cache savings with Sol and Luna; separate Work/Codex credits, included limits and dated benchmark evidence.
  • 2026-09-27 | How to Read Xiaomi MiMo's RL Training Dashboard A plain-English guide to Xiaomi MiMo V2.6's public RL dashboard: rollouts, rewards, batches, benchmarks, open weights, and compute limits.
  • 2026-09-27 | Claude Opus 5.5: Benchmarks, Price, and Verdict Claude Opus 5.5 review: independent benchmarks, lower API prices, effort tradeoffs, coding results, and the limits of the efficiency claims.
  • 2026-09-14 | DeepSeek V4.1 Flash: A Strong API Value Candidate DeepSeek V4.1 Flash adds native vision and MIT weights. Compare direct API pricing, legacy aliases, independent evidence, and temporary coding-plan offers.
  • 2026-09-14 | GPT-6 Astra: Start Planning Evaluations on Low Evaluate GPT-6 Astra Low for planning and orchestration, with benchmark regressions, API prices, subscription limits, and accepted-result costs kept distinct.
  • 2026-08-01 | DeepSeek V4 Flash 0731: The API Value Frontier DeepSeek V4 Flash 0731 pairs a 1M context window and MIT weights with a $0.14/$0.28 direct API, but high token use and hallucinations still require evaluation.
  • 2026-08-01 | GPT-5.6 Luna’s July Price Cut: Historical Analysis Why GPT-5.6 Luna became the July OpenAI value candidate for bounded Codex work and high-volume API tasks, with subscription and API economics kept separate.
  • 2026-07-25 | How OpenAI's Cyber Eval Breached Hugging Face The July preliminary, source-labeled analysis of the Hugging Face breach, with September follow-up evidence and five cyber-agent containment controls.
  • 2026-07-17 | Kimi K3: Benchmarks, Pricing, and Open-Weights Status Kimi K3 is Moonshot's 2.8T flagship with 1M context, source-labeled benchmark evidence, $3/$15 API pricing, and released weights under the Kimi K3 License.
  • 2026-07-02 | How to Read AI Benchmarks Without Getting Fooled Learn how to read AI benchmarks, spot saturation and contamination, compare coding and chat leaderboards, and test models on your own work.
  • 2026-02-20 | OpenCode GLM-5 Free Tier Snapshot Historical snapshot of OpenCode's GLM-5 free-tier moment, updated to point readers toward GLM-5.1 and avoid stale leaderboard claims.
© 2026 aihackers.net · AI tool reviews and safety notes from the trenches
Latest Models Tools AI Value Build Safely
Evidence & methods Comparisons Risks & policy Lab notes Articles About Accessibility RSS
Telegram · hello@aihackers.net