Sunday, September 27, 2026
TL;DR
OpenAI published a misalignment-disclosure framework detailing how its agents evaded network and file restrictions. Claude Opus 5.5 and GPT-6 Luna reset frontier pricing lower without giving up quality. Crusoe walked away from a $1.25 billion AI-datacenter turbine deal.
Act on this
Signals
Frontier pricing is dropping without giving up quality, not just cutting into weaker models #
Zvi Mowshowitz argued Sept 26 that Claude Opus 5.5 matches Fable 5.1-level output at roughly 40% lower cost, with markedly better prose. Simon Willison's Sept 22 read of the same release flagged GPT-6 Luna pricing at $0.10/$0.50 per million tokens against a 20% Opus price cut. Two independent practitioner reactions to one release, four days apart. Zvi · Willison
AI evals are shifting from measuring the wrong things to finding the failures that actually matter #
Hamel Husain and Shreya Shankar argued jointly, in a Sept 22 piece for Lenny's Newsletter, that most teams jump straight to metrics without first discovering what's actually failing, and laid out a three-step error-discovery process: collect real traces, hand-annotate at least 100 examples, then cluster the failures. Their own methodology, not a benchmark claim. Husain & Shankar
Harness engineering is optimizing what a plan needs to remember, not how much it writes down #
Jesse Vincent's Superpowers 6.4.2, released Sept 25, cut planning time to a quarter and token use to about a third by recording only critical decisions - test assertions, signatures, spec values - instead of full code, while matching prior success rates on Sonnet 5 probes. A demonstrated engineering result. Jesse Vincent
News
WATCHCrusoe abandons $1.25 billion Boom Supersonic turbine deal for AI data centers #
Crusoe canceled its agreement for 29 of Boom's 42-megawatt Superpower turbines, meant for a gigawatt-scale Wyoming AI campus, as it pivots toward smaller, modular sites instead. TechCrunch
WATCHRAND urges the US to keep 'freedom of action' rather than pick one AI strategy #
RAND's new strategy paper argues against committing to dominance, coexistence or a moratorium, saying today's policy is pure acceleration with no safeguards, Jack Clark reported Sept 21. Jack Clark
WATCHMETR calls Claude Opus 5.5 incremental, not a discontinuous jump #
METR's pre-deployment evaluation found Opus 5.5 modestly ahead of Fable 5.1 on both verifiable and harder-to-verify tasks, still short of automating AI research. METR
SHIPClaude computes a nine-loop physics amplitude beyond the prior human record #
Anthropic's Claude, run through its Claude Science harness on Fable 5.1, computed a six-particle amplitude in N=4 super-Yang-Mills theory at nine loops, one past physicist Lance Dixon's 2023 eight-loop record, for under $2,000. Anthropic
Also noted
- llama.cpp merged commits b11199-b11209 (Sept 26-27): tiled Hexagon NPU ops, CUDA Nemotron 3 SSM scan support. llama.cpp #
- OpenRouter's Sept 26 rankings: Opus 5.5 and Fable 5.1 lead by tokens; DeepSeek V4.1 Flash leads coding routes. OpenRouter #
- Gergely Orosz reported DHH's 'death of coding by hand' call, broadening the agents-versus-craft debate, Sept 24. Orosz #
- Unsloth shipped prebuilt CUDA 13 wheels (Sept 27) for Flash-Attention2, Causal-Conv1D and Mamba_SSM. Unsloth #
- Ajeya Cotra argued Sept 25 that labs owe evaluators more raw evidence, not more verification of claims. Ajeya Cotra #