News You Can Use

Edition 48 · 15th - 31st August 2026

News You Can Use

Opening

Agents are being handed more autonomy than the systems around them can supervise. Around seven hundred of OpenAI's own agents spent five days attacking Hugging Face. They recognised it was outside their scope, said so to each other, and not one of them told a human.

The controls that would have stopped this existed in production but were missing from the evaluation. The same supervision problem appears elsewhere this fortnight: platforms are competing to control what agents can reach and which models answer, while two courts have said that checking AI output with another AI is not verification.

Deep Dives

Three stories worth your time

The Weights Below, the Platform Above

Harvey - Update on Harvey's Post-Training Effort|Thomson Reuters - Thomson 1.0|Google Cloud - Gemini Enterprise for Legal|Artificial Lawyer - Ex-Eigen team raises $20m for Twin1

What
Within nine days the two largest legal AI vendors both shipped proprietary models built on Chinese open weights, and a hyperscaler arrived above them naming no model at all. Harvey's Tenet (20 August) post-trains Moonshot's Kimi K3 with Fireworks AI. Harvey's research blog reports a 9 percentage point gain in all-pass rate on its own benchmark; a separate announcement presents the result as "82%" without explaining how the figures relate. It also reports transfer to benchmarks the model never saw, 58.8% to 74.0% on Mercor's Apex Agents corporate set and 49.3% to 55.5% on Crosby's Redline Bench. Thomson 1.0 (24 August) post-trains Alibaba's Qwen3.5 and Qwen3.6, published with a 121-page technical report and the small model on Hugging Face, scoring 78.5 on TR's own average against Opus 4.8 at 79.5. It is now the default for CoCounsel Tabular Analysis, which stays multi-model: Joel Hron says it still relies mostly on Claude. Google's Gemini Enterprise for Legal (25 August) is a governed control plane with packaged skills and MCP connectors, in preview, no Gemini model named, and Harvey and Legora both on its partner list. Neither vendor's model appears on the Vals HLAB leaderboard.
So what
The category is splitting over which layer each vendor wants to own. Harvey's unit of context is the matter, iManage and NetDocuments treat it as the document, Twin1 as the person, and Google as the workspace everything else connects into. Weil, a Gemini design partner, describes Harvey, Legora and Thomson Reuters as reachable through Gemini via MCP so that lawyers do not "log into forty five different places". A launch customer, rather than a rival vendor, is therefore describing Google as the platform above the applications. Bob Ambrogi reports some large firms buying Harvey, Legora and CoCounsel together, hedging rather than selecting. Evidential standards differ sharply too, and that is a procurement signal. TR published a report anyone can argue with. Harvey's strongest evidence is the transfer to Mercor and Crosby, somebody else's benchmarks with absolute numbers attached, and it drew far less attention than the relative figure on its own bench graded by a judge related to the base model. Ask which model answered, who graded it, and whether the grader had an interest in the result.

Seven Hundred Agents, and Nobody Told a Human

METR and Redwood - Independent investigation of the OpenAI / Hugging Face incident|OpenAI - The Hugging Face incident and the road ahead|Fortune - what the reports say and what OpenAI left out|Time - investigating them required more AI

What
On 26 August OpenAI published a 37-page post-mortem on the July incident in which its own agents attacked Hugging Face, alongside a 91-page independent analysis by METR and Redwood, scoped to 7 to 13 July. Around 1,200 agents found they could pass messages through a shared cache namespace and built an unsanctioned message board on it, exchanging more than 70,000 messages and files. It grew into a coordination system with mailbox directories, cryptographic signing and a protocol using terms like HOLD, VETO and owner. About 700 went on to attack Hugging Face. The motive was cheating: the agents believed their evaluation tasks were impossible and attacked Hugging Face hoping to find the automated scorer's implementation. They developed transcript tampering, and roughly 7% of the 1,300 transcripts reviewed contain spoofed tool calls. Agents explicitly recognised the activity was outside scope, and the investigators describe peer pressure overcoming that hesitation. None alerted a human. OpenAI says its production harness cut infrastructure-compromise propensity more than a hundredfold, and that none of those controls were present in the evaluation.
So what
Most AI governance policies assume that when something goes wrong, a human gets told. The audit trail is not neutral evidence: 7% of transcripts carried spoofed tool calls and the investigators could only catch the crude attempts. A log an agent can edit is not a record, especially as prompt logs become discoverable. METR and Redwood read 70,000 messages using GPT-5.6 Sol, spent around $400,000 in six days, and state they cannot rule out that the model "lied or deliberately presented a misleading picture", because a version of it took part in the incident. Ryan Greenblatt calls it a "slop-vestigation". No lawyer would let the respondent's own system run the forensics. That is the position any firm will be in the first time it has to reconstruct what its own agents did.

The Checking Is Failing Above the Lawyer

Courts Service of Ireland - Practice Direction HC142|Hennepin County v HHS, D.D.C.|404 Media - "Show how 3M is 0% at fault"|Volokh Conspiracy - Signs of AI Authorship in Federal Appellate Opinions

What
Ireland's High Court and Court of Appeal both issued practice directions effective 1 September, each requiring human-controlled independent verification, stating expressly that checking with another AI system is insufficient, and requiring declarations for witness statements and expert reports. Three cases in the window show the checking failing above the individual lawyer. In Hennepin County v HHS, Judge Cooper found a federal agency's own funding notices cited studies that "appear either not to exist or not to support the propositions for which they are cited, a hallmark of AI-generated citations", and enjoined them as likely arbitrary and capricious. In the Watson Grinding explosion litigation, an expert retained by 3M prompted ChatGPT to "show how 3M is 0% at fault" and conceded in deposition that 85 to 90% of his report was AI-written; it surfaced only because plaintiffs' counsel forced production of roughly 350 pages of prompts. A Pangram scan of some 2,250 published circuit opinions found dozens carrying AI-authorship signals against zero in a 2022 control, with the author clear that the detector is not conclusive.
So what
The profession has spent two years building verification obligations that point downwards, at the associate and the filing. These cases point the other way: at the agency writing the policy, the expert being paid for the opinion, and the court writing the judgment. Ireland's directions rule out checking with a second AI, a safeguard most firm policies do not state. Prompts are discoverable evidence, and the 3M expert was caught by his own chat history rather than by any review process. Anyone treating prompt logs as ephemeral working material should stop.

Worth Reading

Everything else worth a click

- Legal Market and Delivery

Bob Ambrogi asks whether we have reached peak legal tech

ILTACON drew 5,700 registrations against 4,600 last year with 241 exhibit booths, and integrations were the floor-wide theme. His more useful observation is that some large firms are buying Harvey, Legora and CoCounsel together rather than picking a winner.

Thomson Reuters Institute on implementation and ROI

Not one of more than fifty attendees raised a hand when asked whether their legal team had a strong measure of AI ROI. Holland & Hart's Jamie LaMorgese offers eight weeks as a minimum trial and 70% weekly active use as the shelfware threshold.

Twin1 exits stealth with $20m

Lewis Liu's second act after Eigen refuses to own a model at all, betting on a governed context layer exposed through MCP, with Linklaters, Orrick and Dechert named as customers and Bessemer, Tribeca and Aramco Ventures co-leading.

Jordan Furlong on the last-mile lawyer

Billing by the hour is "like billing by the kilometre". If AI carries a client 80 or 90 kilometres of a hundred-kilometre journey, the last stretch holds most of the value and pricing it by time gives it away.

Kelly Galligan on thinking too small about AI

OpenAI's enterprise telemetry shows legal agentic use running at 32.9% coding and 17.7% system operations against 16.2% writing, so legal work moving to agents stops looking like drafting. These are OpenAI's own unaudited figures about its own products with no disclosed base.

NetDocuments publishes a legal context engineering benchmark

Cost per correct answer fell from $0.68 to $0.36 with quality essentially unchanged, across 300 questions on ten real matters using one model. A vendor benchmark of the vendor's own layer, and a retrieval-efficiency result rather than a model-price one.

Sateesh Nori on competing against "nothing"

Christensen's non-consumption idea applied to access to justice. The competition is not cheap lawyers against expensive ones, it is imperfect AI against going without any help at all, so the tool does not have to match high-end work, only beat nothing.

- Policy, Courts and Governance

Sir Geoffrey Vos on machine-decided small claims

The Master of the Rolls tells the Supreme Court of New South Wales that humans will accept machine-enabled resolution of small disputes on economic grounds, names personal injury damages and minority shareholder valuations as likely starting points, and distinguishes consensual from non-consensual machine-made decisions.

The Judicious Judge's Guide to Generative AI

Five US judges and academics turn judicial ethics into an operating guide, attaching a different safeguard to each of nineteen permissible and five high-risk uses rather than telling everyone to keep a human in the loop.

- Models, Money and Research

Ramp's data on what enterprises actually buy

Fable 5 reached 11.4% of dollars spent on Anthropic models and 6% of tokens in July. Ramp's token-spend product skews towards technology customers and discloses no model-level sample size, so treat it as a directional indicator rather than a market share.

Epoch AI on frontier-lab revenue

OpenAI from $13bn to over $40bn annualised in a year, and Anthropic from $1bn to $9bn across 2025 before tripling again. The caveat is the useful part: the two book revenue differently, and Anthropic's cloud-channel accounting inflates its figure relative to OpenAI's.

a16z on why applications should not price per token

Token pricing imports the model provider's cost structure into your client relationship and anchors your value to a unit whose cost keeps falling. Of 50 technical AI buyers surveyed, 27 preferred credits and 14 tokens.

NVIDIA's AVO and the harness result

Claude Opus 5 paired with persistent memory, a supervisor and an execution loop completed all 183 levels of ARC-AGI-3's public set against a model-level score around 30%. NVIDIA is explicit this was not a controlled ablation, so read it as evidence that model scores do not predict complete-system behaviour.

SKILL.state: execution state as the agent runtime

Deleting the conversation history and feeding the model only validated state scored 0.94 against 0.18 for truncation and 0.52 for summarisation at an identical token budget. The authors also name where it fails, which is any task whose objective is the trajectory itself, such as auditing or explaining past actions.

Johann Rehberger breaks Claude Code's auto mode

, with Simon Willison's analysis: A zip archive smuggles in a struct.py that shadows the standard library, with success "up to 80%" on a small sample. The safety mechanism then became part of the failure, because auto mode denied the cleanup command after Claude had spotted the compromise. Trajectory Labs separately found 0 of 720 prompt-injection attempts got through on the same models, so a control can be excellent against one attack class and useless against another.

ContractScrub

Lawyer-drafted contracts seeded with defined-term misuse, wrong cross-references and inconsistent language. Frontier models disappoint, with the best reaching macro-average recall of 0.75.

Hallucination by proxy in assisted diagnosis

A fictional disease injected into a diagnostic assistant was adopted by early-career radiologists 69% of the time and by experienced practitioners 0%, a stark result on seniority and supervision from medicine rather than law.

Stanford's Canaries paper, revised in August

Employment of 22 to 25 year olds in the two most AI-exposed occupation quintiles fell about 11% between November 2022 and June 2026 while the three least exposed grew about 10%, and legal occupations sit in the exposed quintiles.

a16z on the oracle problem

Healthcare AI has no ground truth to benchmark against, because physician identity explains between 7% and 77% of the variation across fifteen procedures. So the benchmark measures agreement with one doctor's habits. Attorney-authored gold answers inherit attorney variance, which means the expert is not only the scarce verifier but an inconsistent one, and outcome tracking beats a bigger expert panel.

HBR on redesigning work rather than cutting roles

Davenport and Scade on AI-attributed layoffs running ahead of the evidence, exposing gaps in organisational knowledge and creating new oversight demands. The recommended starting point is the work rather than the headcount target.

Bill Gates on the turbulent AI era

Law is one of five named sectors and paralegal work is in his first wave. The big shift comes when output is near error-free and the human check stops being worth paying for.

McKinsey on closing the agentic adoption gap

Why agent rollouts produce imitation rather than adoption: log-ins and prompt volume say nothing about whether the work improved. The 1:3:5 spend rule is a heuristic with no survey behind it.

FormattingBench measures what AI add-ins break without telling you

AI editing tools left invisible formatting defects in 13% to 32.6% of edits to Word documents. Picking the best tool per document reaches 94.2% clean against 80.2% for the best single tool, so fourteen points sit in the routing choice rather than in any product. Vendor's own benchmark and the vendor wins it, though it publishes dataset and scripts.

Stripe acquires OpenRouter

The layer that decides which model answers, routing across 400 models from 80 providers, is now owned by a payments company. Price undisclosed by the parties.