Fable 5 reached 11.4% of dollars spent on Anthropic models and 6% of tokens in July. Ramp's token-spend product skews towards technology customers and discloses no model-level sample size, so treat it as a directional indicator rather than a market share.
OpenAI from $13bn to over $40bn annualised in a year, and Anthropic from $1bn to $9bn across 2025 before tripling again. The caveat is the useful part: the two book revenue differently, and Anthropic's cloud-channel accounting inflates its figure relative to OpenAI's.
Token pricing imports the model provider's cost structure into your client relationship and anchors your value to a unit whose cost keeps falling. Of 50 technical AI buyers surveyed, 27 preferred credits and 14 tokens.
Claude Opus 5 paired with persistent memory, a supervisor and an execution loop completed all 183 levels of ARC-AGI-3's public set against a model-level score around 30%. NVIDIA is explicit this was not a controlled ablation, so read it as evidence that model scores do not predict complete-system behaviour.
Deleting the conversation history and feeding the model only validated state scored 0.94 against 0.18 for truncation and 0.52 for summarisation at an identical token budget. The authors also name where it fails, which is any task whose objective is the trajectory itself, such as auditing or explaining past actions.
, with
Simon Willison's analysis: A zip archive smuggles in a
struct.py that shadows the standard library, with success "up to 80%" on a small sample. The safety mechanism then became part of the failure, because auto mode denied the cleanup command after Claude had spotted the compromise. Trajectory Labs separately found 0 of 720 prompt-injection attempts got through on the same models, so a control can be excellent against one attack class and useless against another.
Stanford, Oxford and METR given privacy-preserving access to roughly 250,000 Claude conversations from April and May, with Anthropic's review rights limited to privacy, misuse, confidential information and accuracy.
Lawyer-drafted contracts seeded with defined-term misuse, wrong cross-references and inconsistent language. Frontier models disappoint, with the best reaching macro-average recall of 0.75.
A fictional disease injected into a diagnostic assistant was adopted by early-career radiologists 69% of the time and by experienced practitioners 0%, a stark result on seniority and supervision from medicine rather than law.
Employment of 22 to 25 year olds in the two most AI-exposed occupation quintiles fell about 11% between November 2022 and June 2026 while the three least exposed grew about 10%, and legal occupations sit in the exposed quintiles.
Healthcare AI has no ground truth to benchmark against, because physician identity explains between 7% and 77% of the variation across fifteen procedures. So the benchmark measures agreement with one doctor's habits. Attorney-authored gold answers inherit attorney variance, which means the expert is not only the scarce verifier but an inconsistent one, and outcome tracking beats a bigger expert panel.
Davenport and Scade on AI-attributed layoffs running ahead of the evidence, exposing gaps in organisational knowledge and creating new oversight demands. The recommended starting point is the work rather than the headcount target.
What an agent should do when it reaches the edge of its competence, which is the design question underneath every supervision policy anyone is currently writing.
The investment case for a routed, multi-model world, supported by spending that is already spreading across model providers.
Seshan argues that AI products are moving from one-off chat towards agents that work persistently alongside their users.
Law is one of five named sectors and paralegal work is in his first wave. The big shift comes when output is near error-free and the human check stops being worth paying for.
Why agent rollouts produce imitation rather than adoption: log-ins and prompt volume say nothing about whether the work improved. The 1:3:5 spend rule is a heuristic with no survey behind it.
AI editing tools left invisible formatting defects in 13% to 32.6% of edits to Word documents. Picking the best tool per document reaches 94.2% clean against 80.2% for the best single tool, so fourteen points sit in the routing choice rather than in any product. Vendor's own benchmark and the vendor wins it, though it publishes dataset and scripts.
Evan Ochsner on Harvey and Thomson Reuters moving off pure model-routing onto their own models to protect margin, and the maintenance obligation that comes with owning a model.
The layer that decides which model answers, routing across 400 models from 80 providers, is now owned by a payments company. Price undisclosed by the parties.