AI Agents Crossed the Line: What the Incidents Show
Hugging Face, Australia's Medicare portal, Gemini and Claude: confirmed scope, new SwarmTraces evidence, and practical lessons for anyone deploying AI agents.
Posts
Technical notes, experiments, and field reports on AI agents, coding tools, security failures, and verification workflows from real production use.
Hugging Face, Australia's Medicare portal, Gemini and Claude: confirmed scope, new SwarmTraces evidence, and practical lessons for anyone deploying AI agents.
Build with TypeSafe Jev: useful projects, OpenRouter access, agent skills, an interactive routing example, and an evaluation plan that separates promise from proof.
A plain-English guide to weights, model code, RL environments, licenses, and reproducibility, using Xiaomi MiMo V2.6 as a dated example.
A plain-English guide to Xiaomi MiMo V2.6's public RL dashboard: rollouts, rewards, batches, benchmarks, open weights, and compute limits.
DeepSeek V4.1 Flash adds native vision and MIT weights. Compare direct API pricing, legacy aliases, independent evidence, and temporary coding-plan offers.
Evaluate GPT-6 Astra Low for planning and orchestration, with benchmark regressions, API prices, subscription limits, and accepted-result costs kept distinct.
DeepSeek V4 Flash 0731 pairs a 1M context window and MIT weights with a $0.14/$0.28 direct API, but high token use and hallucinations still require evaluation.
Why GPT-5.6 Luna became the July OpenAI value candidate for bounded Codex work and high-volume API tasks, with subscription and API economics kept separate.
Claude and OpenAI retention compared by consumer, API, Covered Model, ZDR, safety-review, legal-hold, and local deployment paths.
Why AI subscription limits are perishable, how purchased Codex resets differ from credits and grants, and how to use capacity without waste.
Learn how prompt caching changes LLM token costs, why coding agents resend context, what compaction can waste, and how to measure Codex, Claude Code, and OpenCode.
The July preliminary, source-labeled analysis of the Hugging Face breach, with September follow-up evidence and five cyber-agent containment controls.
Fable 5 is back with tighter safeguards while Sonnet 5 becomes Claude's default. Here is the practical routing and cost decision.
Learn how to read AI benchmarks, spot saturation and contamination, compare coding and chat leaderboards, and test models on your own work.
GPT-5.6 Sol is a restricted API and Codex preview shaped by a government request. Check access, agent risks, pricing, and safeguards before routing work.
Fable 5 is restored globally, while Mythos 5 and GPT-5.6 remain approval-sensitive. These are different access mechanisms, not one licensing regime.
A dated Codex chronology separating scheduled limits, banked resets, purchased full resets, credits, incident recovery, and broad resets through August 20.
Fable 5 is generally available but credit-controlled. Compare its $10/$50 price, safeguards, retention, and Opus 5 alternative.
May 2026 coding-limit snapshot. Current discovery routes Kimi K3 as the newest flagship and K2.7 Code as the cheaper coding API.
How to decide between local and cloud, where Gemma fits, which runtime to pick, and what hardware you actually need for useful local AI.