News You Can Use

Edition 47 · 1st - 14th August 2026

News You Can Use

Opening

“AI use” has become too blunt a category to tell us very much. The useful questions are what work we delegate, whether the system has the context it needs, and how we prove that someone remained accountable for the result.

This edition looks at evidence across these three areas. The pattern is fairly consistent: AI works best when it sits inside a deliberate process, and worst when simply including it is seen as enough.

Deep Dives

Three stories worth your time

Delegation Is the Variable

University of Minnesota - Artificial Intelligence and Human Legal Reasoning|Shen and Tamkin - How AI Impacts Skill Formation|Abdulhai et al. - How LLMs Distort Our Written Language

What
In a preregistered Minnesota trial involving 91 upper-level law students, using Gemini 2.5 Pro to synthesise an unfamiliar legal doctrine increased the quality of the initial work by 59.8%. When AI was removed for the next task, the group showed no immediate comprehension loss. Its advantage on a later application memo disappeared once the quality of the earlier synthesis was controlled for: AI had improved the material they were reasoning from, rather than their underlying reasoning. When everyone later used AI to revise their own memo, weaker work generally improved while some stronger work deteriorated, although regression to the mean makes the size of that decline unsafe to use. In a separate preregistered study of 52 professional developers, AI users scored 4.15 points lower on a 27-point comprehension test and received no speed benefit. Those who generated code and asked one follow-up question scored 86%; those who stopped after generating it scored 39%. A third study found no measurable distortion when AI was used as an information tool, but wholesale delegation made writing more homogeneous and could alter its argument.
So what
Whether AI helps or harms depends heavily on how much of the thinking is delegated. It can improve the artefact that carries work into the next stage, while also preventing the user from acquiring the knowledge they would have picked up by doing it themselves. This explains why genuine practitioner enthusiasm and disappointing organisation-wide results can coexist. Deployment should therefore be designed around behaviours, not usage totals: break work into defined components, require users to interrogate and explain outputs, and avoid mandatory AI revision at the end of work that was already good. Asking whether somebody used AI tells us very little. Asking whether they could defend the resulting work in a demanding conversation, without going back to the model, is a much better test.

The Matter Has to Survive the Journey

Harvey Labs - Law Firm Knowledge benchmark|LawSites - Agent Handoff Protocol|Agent Handoff Protocol

What
Harvey and EngramLab created Calderwood & Harkness, a synthetic law firm containing 46 clients, 266 matters and 9,288 files across 15 practice areas. Harvey then tested GPT-5.6 Sol and Claude Opus 4.8 on 250 tasks against that shared institutional record. Both models generally found the core information but struggled to know when they had found everything, particularly where a task required several separate facts from across the corpus. Harvey’s published baseline said the models satisfied around half the grading criteria, although the rubric was revised four days later, so the qualitative failure is more durable than the percentage. Separately, DeepJudge published the draft Agent Handoff Protocol, with Harvey planning beta support and Thomson Reuters intending to support it in CoCounsel. AHP moves a user and an approved package of task context between complete agent products: an objective, selected messages, files and a thread identifier. System prompts, hidden reasoning, credentials and execution privileges stay behind.
So what
Legal work does not begin and end inside one prompt or one product. A matter accumulates facts, decisions, permissions and work product over months, then crosses between research, drafting, review and specialist tools. Re-uploading documents and re-explaining the task at every boundary loses information and creates inconsistent copies. Giving every product permanent access to everything creates a different problem. AHP is an early proposal rather than an adopted standard, but it identifies the right unit of portability: selected matter context under the user’s control. Harvey’s benchmark shows why that context also needs structure. Search alone cannot tell an agent whether its answer is complete. Firms will need permissioned indexes, summaries and memory that persist across tasks, alongside auditable rules governing what moves between systems. The context layer is becoming both the source of useful performance and the next source of switching costs.

A Watermark Is Not a Confession

Courts Service of Ireland - Practice Direction HC 142|Anthropic - How Claude marks AI-generated content|Daniel Miessler and Kai Magnus - Where an AI Watermark Can Hide in Plain Text

What
Ireland’s High Court issued Practice Direction HC 142 on the responsible use of generative AI in court documents, effective from 1 September. It applies to pleadings, submissions, affidavits, witness statements and expert reports in new and existing civil proceedings. Responsibility remains with the party, lawyer, witness or expert regardless of AI use, and witnesses must declare that AI was not used to generate substantive content in statements or affidavits. Formatting and spellchecking are permitted. Meanwhile, Anthropic has begun applying invisible text watermarking and C2PA provenance metadata to Claude outputs globally as part of its response to the EU AI Act’s transparency rules. Anthropic has not disclosed its technical implementation. An informed reconstruction by Kai Magnus and Daniel Miessler suggests durable text watermarking is likely to rely partly on statistical patterns in model word choice. Such a signal can survive copying but weakens as the text is rewritten or paraphrased.
So what
Detection cannot settle the question a court declaration asks. A human-written document passed through Claude for proofreading may carry a mark, while heavily rewritten AI-generated text may not retain enough of one to detect. The mark can indicate that Claude processed the text at some stage; it cannot by itself establish who authored the substance or how much judgement the human exercised. That distinction reaches beyond court filings into privilege reviews, disclosure exercises and internal investigations. Organisations should record AI use when it happens rather than trying to reconstruct authorship from the finished document: what went into the system, what came back, what changed, who checked it and who approved it. Technical provenance can support that record, but responsibility still rests with the person signing their name.

Worth Reading

Everything else worth a click

- Legal Market and Delivery

Align Research goes generally available

An AI research product that retrieves relevant cases without generating an answer, a useful counter-trend to systems that hide retrieval and synthesis behind one confident response.

Harbor launches Deploy

The legal consultancy is embedding forward-deployed engineers inside client organisations, an explicit response to the deployment teams now being built by OpenAI, Anthropic and Microsoft.

Thomson Reuters partners with Laurel on AI measurement

Time and activity data will be used to measure the effect of AI inside professional workflows, although Laurel's claims of up to 30 minutes recovered per lawyer per day and 4-11% revenue uplift remain company-reported.

DISCO moves beyond e-discovery

Its Unified Litigation Solution combines matter facts with external case law and is in pilot with five firms, with general availability planned for early 2027.

Relativity announces claiR

Conversational analysis across a full RelativityOne matter, including metadata and relationships, is in Advanced Access with three firms; general availability is expected in early 2027.

- Policy, Courts and Governance

What 100 recent AI cases say about sanctions

Adam Feldman's review finds that consequences are driven more by what lawyers do after discovery than by the initial error, with prompt admission generally associated with a better outcome.

The AI Growth Lab opens to legal innovators

The SRA, Legal Services Board, Council for Licensed Conveyancers and ICO are offering a single regulatory entry point for testing, with applications open until 27 September.

- Models, Agents and Risk

Anthropic tests emerging multi-agent systems

More capable individual agents did not reliably produce better groups, with experiments showing duplicated work, poor resource allocation, conformity and occasional collusive or destructive behaviour.

Open weights are not open source

Gary Marcus explains why releasing trained parameters without training data, preprocessing or algorithms does not provide the transparency that procurement teams often assume the word "open" carries.

OpenAI previews Ultrafast

GPT-5.6 Sol runs on Cerebras at a claimed maximum of 750 output tokens per second, turning frontier-model latency into a service tier rather than a reason to choose a smaller model; price and general availability remain undisclosed.

- Capability and People

How organisations use ChatGPT

OpenAI's working paper covers 1,764 organisations and finds early-career workers send eight to nine more weekly messages than the average active user in the same firm, while executives use it less and skew towards reading-shaped tasks.

What OpenAI calls frontier firms do differently

The gap in output tokens per active user between the top 10% and the middle widened from 2.6x in January to 8.3x in June, although token volume remains a poor proxy for value and every figure concerns OpenAI products.

Anthropic reviews worker-retraining evidence

Across 56 randomised US studies plus European evidence, training offers produced modest average employment and earnings gains, with stronger sector-based programmes difficult to replicate consistently.

Josh Kubicki publishes a legal-AI glossary

A practical reset for buyers whose suppliers use the same words for different things, ending on the useful diagnostic question: whose judgement is in this output, and can I see it?