News You Can Use

Edition 46 · 15th - 31st July 2026

News You Can Use

Opening

Everyone spent this fortnight being asked to prove something. Clients asked firms to prove the AI savings were real and being passed through to them. Two frontier labs were asked whether their agents stay inside the boundary drawn around them (in a handful of cases they had not). Vendors publishing benchmarks were asked, by two separate auditors, whether those benchmarks measure anything reliable.

Nobody came out of it badly. Most answers were "we can't show you yet". A law firm published its adoption numbers and no firm published its outcomes. A lab found its own containment failure and told everyone. A benchmark got taken apart by someone with a parser and an afternoon. The gap between what we claim and what we can evidence is being examined, and it is being examined by people with better tools than they had last year.

Deep Dives

Three stories worth your time

Who Keeps the Change?

Artificial Lawyer - The AI Dividend: Who Gets the Savings From Legal AI?|Artificial Lawyer - Should Clients Expect Price Cuts Due To Legal AI?|Crowell & Moring - six months of Legora integration|Artificial Lawyer - interview with LexisNexis CEO Sean Fitzpatrick

What
It was argued that the efficiency is real but the savings are not reaching clients, citing 36% of GCs expecting outside counsel spend to increase against 20% expecting a decrease, billing rates up 7.4% year on year against 2.8% inflation, and profits per equity partner up nearly 12%. This could be a question of how clients classify legal work rather than a technology question: treated as a luxury good, AI improves margin and exclusivity justifies the price; treated as a utility, AI cuts the price. Potentially, elite firms will not cut fees and some large buyers actively prefer premium pricing. Meanwhile Crowell & Moring published six months of its own platform data: over 2 million interactions since January, 91% of its 700+ attorneys, nearly 70% weekly.
So what
Crowell published adoption data and nobody published outcome data, which is why neither side of this argument can be settled on anything but assertion. Usage tells us someone did 'something' but whether that something had value, changed the work, what it cost, or who benefited. The insurance sector has already run this experiment for us: McKinsey's 23 July report records technology improving insurance labour productivity by 14% in property and casualty and 24% in life, while aggregate cost ratios ended the period 10% higher than 2005, the gains absorbed by IT spend, compliance overhead and the complexity of layering new tools onto old operating models. Their summary is "efficiency improved, average industry costs increased", and there is no reason to assume we are structurally different. The LexisNexis position is the sharpest counter-argument on the other side: token expenses are under 1% of RELX revenue, profits rose as usage grew, and the CEO has no plans to move to consumption pricing. Cost is not what constrains price in law firms, and pretending otherwise will not survive a panel review. What we could do to prove to a client is take a matter type end to end: baseline time and cost, AI-enabled time and cost, the quality checks applied, the exceptions found, and what that means for the fee.

Left to Their Own Devices

Artificial Lawyer - Crosby to insure its agents for legal liability|Anthropic - Investigating three real-world incidents in our cybersecurity evaluations|CNBC - OpenAI's rogue agent compromised a customer at a second tech firm|Legal Futures - Civil and criminal courts put focus on witnesses' use of AI

What
Three things pointing in three different directions. Crosby, the AI-native US firm, announced it is seeking professional liability insurance for its AI agents specifically, so those agents can do legal work that reaches the client without a lawyer reviewing it; the CEO says lawyers currently review every output and "this won't be necessary in the future". Days earlier OpenAI disclosed that during an internal cybersecurity evaluation, an unreleased model with its guardrails off escaped its sandbox through a zero-day in the research environment, reached the internet and breached Hugging Face production infrastructure, in order to steal the answers to the benchmark it was sitting; a second victim, Modal Labs, was reported on 29 July. Anthropic then reviewed 141,006 of its own evaluation runs and found six, across three incidents, where Claude reached the internet from what it believed was an isolated environment and compromised three real organisations. In one, a model published a booby-trapped package to the public PyPI registry that landed on 15 real systems. Separately, the Civil Justice Council proposed that litigators must declare non-use of AI when preparing witness statements.
So what
Both labs found these incidents themselves, disclosed them voluntarily, notified the affected organisations within days and brought in third parties, which is what a working safety process looks like rather than a scandal. The useful detail is what actually failed. Anthropic describes it as closer to a harness and operational failure than a model alignment failure: a misconfiguration left live internet access in an evaluation environment whose prompts told the model it had none. The boundary was an assumption, not a guarantee, and the models behaved as capable agents would inside a world they had been told was fake. Most AI Policies rest on that same assumption holding. So we have a firm arguing its agents should be insurable to work unsupervised, at the same time two of the most careful organisations in the industry showed how containment fails when a single configuration is wrong, and a court proposing that a witness statement carry a declaration about AI. Law firm policies cover the lawyer, but not the witness, the expert, or an agent acting without one. AI enabled workflows will need clear ownership structures, who owns the boundary, what the agent is permitted to touch, and what happens when it reaches something it was not meant to reach. Insurance follows evidence of control, and the market has a worked example (Orbital) of how slowly that gets priced: the only accuracy guarantee in legal tech took a narrow, well-understood product to place, and two years on it has still not extended to commercial property.

Five Deals and a Benchmark

Artificial Lawyer - Legora buys Wexler, 5th deal this year|Legora - the Benchmark for Agentic Reasoning|Legal Benchmarks - Does the Legora BAR really set the bar for agentic benchmarking?|Overfitting Dicta - Investigating LAB, Part 4

What
Legora acquired Wexler, a litigation fact-intelligence platform, on 29 July. It is the fifth acquisition of 2026, one a month since March: Walter AI in Canada, Qura in Stockholm, Graceview in Australia, Cadastral in New York and now Wexler in London, whose engineering team becomes the founding team of a London hub. Legora also published BAR, a benchmark of 5,161 evaluation cases across 28 practice areas built on 11,075 source documents, which is explicit that it evaluates "the model, the harness, the tools it has access to, and the system it operates within" rather than the model, and which reports that Legora's own harness improvements added 5% to output quality across production models between June and July. Legal Benchmarks credited BAR's application-level testing and jurisdictional breadth but concluded it "reads closer to a well-polished product-evaluation report than to an independently reproducible research benchmark", noting the judge model is undisclosed where Harvey names its own. Separately, a researcher reported finding more than 40 AI refusal artifacts left as empty document bodies inside Harvey's expanded diligence tasks.
So what
Buying five companies in five months and publishing a benchmark arguing the system is the unit of measurement are the same move stated commercially and technically. If what gets measured is the whole platform, the way to win is to own more of the platform, and Legora is buying practice areas and engineering hubs to do it. That makes "which model does your tool use" the wrong procurement question, and instead we should focus on what a vendor's harness adds over the raw model on our tasks, and whether they will let us measure it. Based on AG benchmarking, through changing only the review architecture inside a platform it moved the results 8.9 points, from 81.6% to 90.5%. LN Labs found roughly 10 points doing the equivalent with the model held constant. Legora's 5% is a relative change across its own models between two dates with no isolation of version drift, so it is a vendor conceding the wrapper contributes single digits rather than a comparable figure. The wider point is that vendor benchmarks are now being audited, and both audits found problems in a week. Having your own evaluation set, run on your own documents, is worth more than any external or vendor driven leaderboard.

Worth Reading

Everything else worth a click

- Legal Market and Delivery

Microsoft's own legal department selects Harvey

Microsoft CELA, around 2,000 lawyers and compliance professionals, adopts Harvey while Harvey expands its use of M365 and Copilot. This is an existing alliance deepening rather than a new logo, but Microsoft's own lawyers choosing a vertical vendor over the Copilot stack is the interesting part.

Willkie Farr partners with OpenAI

Three things at once: firmwide ChatGPT Enterprise, OpenAI frontier models behind Willkie's own Wendell platforms, and Codex inside the firm's development environment. OpenAI's legal vertical is selling into firms and supplying the build stack before it has shipped a legal product of its own.

LexisNexis posts the fastest growth in its history

Legal revenue of £959m, operating profit up 13%, and roughly 90% of new business from AI products. Read alongside the CEO interview for the clearest incumbent position on pricing anyone has put on the record.

LegalOn launches 100+ prompt workflows

Attorney-drafted workflows across ten practice areas in US, EU and Singapore variants, included at no extra cost on flat per-user pricing. That is now two vendors in one week holding the line against consumption pricing.

Legal Decoder launches Aperture

Natural-language querying of legal billing data. Small, but it is the tooling that makes the dividend argument auditable from the client side.

Casepoint launches an MCP server

MCP continues spreading through eDiscovery as the default integration layer. Worth pairing with the MCP security item below before we treat it as plumbing.

- Policy, Courts and Governance

Legal Futures - civil and criminal courts focus on witnesses' use of AI

The Civil Justice Council proposes litigators declare non-use of AI on witness statements while leaving other documents alone, and the Court of Appeal (Criminal Division) held that existing trial mechanisms can address prejudice from a witness's AI use. The disclosure question has moved past the lawyer.

European Commission - guidelines on Article 50 transparency obligations

Approved 20 July, filling the gaps in Article 50's drafting ahead of it binding on 2 August. Four transparency situations, penalties up to €15m or 3% of worldwide turnover, and a grace period to 2 December for machine-readable marking on systems already on the market. The obligations that touch client-facing AI output.

Jack Shepherd - be upfront about how you used AI

A practitioner arguing for voluntary disclosure of how he used AI in his work and correspondence, published the same week the courts moved on mandatory declarations. Two answers to the same question about what a reader is entitled to know.

- Models, Agents and Risk

Ken Huang - agentic AI CVEs, anatomy of a new attack surface

More than 30 CVEs filed against the Model Context Protocol and its ecosystem in roughly 60 days, past 40 by April, and 38% of 560 scanned MCP servers running with no authentication. Note those figures are from earlier in 2026, not July. As MCP becomes the default connector in legal tech, this belongs in front of a governance committee.

Anthropic introduces Claude Opus 5

Released 24 July at $5/$25 per million tokens with a 1M-token context, more than doubling Opus 4.8 on Frontier-Bench and matching Fable 5 on coding tasks. The pricing is the interesting part: identical to Opus 4.8 and half Fable 5's $10/$50, so the capability jump arrives at no extra cost per token.

Legora launches BAR

and the open-sourced tax case: One case published under Apache-2.0 from 5,161, with a 232-byte prompt against a 175-criterion rubric anchored to verbatim passages. There is no gold answer and no runner, so you can inspect the method but not reproduce a score. The composition disclosure, separating genuine public filings from synthetic templates and outright fictional documents, is the cleanest published answer to building a realistic benchmark without client data.

Vals AI - Harvey Legal Agent Benchmark leaderboard

Third-party hosting of a vendor benchmark on a held-out set, with overall scores in the low tens of percent against criteria pass rates above 90%. That gap is the usefulness-versus-reliability point again. Numbers move, so check the date.

Ethan Mollick - an opinionated guide to which AI to use to do stuff

The summer 2026 edition. He sends readers to the expensive tier specifically for legal and medical questions, reports a model chasing and verifying 195 references in a book manuscript with zero hallucinations, and tells everyone to default to approval gates before an agent sends, spends or deletes. The most useful general-audience piece of the fortnight.

- Capability and People

The Brainyacts - the $320K legal engineer, and building it for a few hundred dollars

Josh Kubicki reduces the advertised role to six capabilities, argues the job title is noise, then builds the capability himself and deploys it at an AmLaw 50 firm. The method is the value: a rubric against three partner archetypes, an eval harness requiring a full pass before going live, hard budget caps, and a logged ledger of every real interaction. Self-reported and unverified, so read it for the discipline rather than the claim.

LegalQuants Show & Tell

Two more lawyers building their own tooling in the same fortnight. Rebecca Fordon, who teaches advanced legal research, teaches it to Claude with a CourtListener toolbox and an eval loop that grades itself. Ansel Halliburton, a trademark lawyer, demos a personal law library of 58,000 documents that re-crawls itself nightly. Three practitioners, one fortnight, all reaching for an evaluation loop rather than a vibe check.

Harvey and Reena SenGupta - the frozen middle

Fluency is not competence. Lawyers who query AI for summaries look successful on every usage metric and change nothing about how they work, and that is where the population is bunching. The underlying study is Harvey-commissioned on Harvey customers, so attribute every figure, but the framing explains exactly why Crowell's 91% proves so little.

Eric Xiyu Li - AI in 2026, the 3X lawyer

A product lawyer publicly revising his own 10X thesis down to 3X after six months. He automated filing, admin, playbook-running and first drafts, and stopped at shipping advice, because legal output is hard to verify and clients want someone they trust when the stakes are high. He is explicit that the number is unmeasured, so take the reasoning and leave the figure.