News You Can Use

Edition 45 · 1st - 14th July 2026

News You Can Use

Opening

The first wave of enterprise AI was about getting capable models in front of people, and getting people to engage with them. The ongoing push now is looking at where AI is used, how we price it, what sensible use is and identifying the line between "should use" and "not suitable". Where AI fails, someone still has to own the failure, and the bill for running it at scale is increasing.

This edition covers a landmark Legal Statement that put AI inside the professional standard of care. Clients started steering work towards whoever can prove an AI-enabled outcome, and away from firms that cannot. And the real token cost of production legal AI became public, which turns model routing and job design into questions a firm has to answer rather than an infrastructure detail. None of this is about whether to adopt AI any more - it's about explaining how you are adopting it responsibly and practically.

Deep Dives

Three stories worth your time

Responsibility to use AI

LawtechUK - Legal Statement on Liability for AI Harms|Legal IT Insider - AI liability clarified|Law Society - landmark Legal Statement|Legal Futures - CPS admits putting hallucinated cases before High Court

What
The UK Jurisdiction Taskforce published its Legal Statement on Liability for AI Harms on 7 July. Its headline conclusion is reassuring: existing English private law is largely capable of allocating liability without a bespoke AI-liability regime. They also stated that inappropriate use of AI, choosing an unsuitable system, inadequate due diligence, or failing to validate hallucinated output may all amount to negligence, and, in the right case, a professional could fall below the standard of care by failing to use AI where a reasonably competent peer would have used it. You do not escape responsibility because a chatbot produced the statement. This is persuasive analysis rather than legislation, but previous UKJT statements have been influential with the English courts. Two days later the CPS admitted that two non-existent authorities had found their way into grounds of opposition and were carried into submissions before the High Court. The CPS said generative AI may have been the immediate source, but the operative cause was a human failure to verify; a review of 78 cases found no similar issue. The judge accepted that AI may be useful or even necessary, while stressing the need for proper oversight.
So what
The useful move here is away from the vague instruction to keep a "human in the loop" and towards a defensible professional method. We need to be able to show why a given system was suitable for the task, what sources were checked, who owned the verification, and how an error would have been escalated. A nominal approval clicked at the end of a workflow is not that, and the UKJT has now given a court the language to say so. LLMs are still not perfect, an example from the Legal Benchmarking leaderboard showed that across 34 drafting and 30 extraction tasks, the leading model reaches only 67.6% drafting reliability, and the model with the highest usefulness score (GPT-5.5) manages just 41.2% reliability, illustrating the gap between polished and dependable output. Law Insider asked 534 AI-adopting transactional lawyers what they most want added to their tools, verification and sources came top at 20% of responses, ahead of everything else. We need to look to replace generic human-oversight language in our own guidance with a task-level verification record: approved system, permitted use, named reviewer, sources checked, exceptions found, and the escalation route. More questions will begin appearing focusing on how you prove that something has been verified, especially from clients with audit powers.

Are in-house teams outpacing their law firms?

Reuters - Norm Ai reaches $1.2 billion valuation|Bloomberg Law - Ford GC says in-house is outpacing firms

What
Norm Ai raised $120 million at a $1.2 billion valuation for a model that pairs software with human-supervised legal and compliance delivery through affiliated Norm Law. The company says its clients represent more than $30 trillion in assets under management, and it is explicitly selling outcomes rather than seats. That supply-side move matches what buyers are saying. Ford GC Steven Croley says the company's internal AI adoption is now outpacing what it sees from external counsel, spans the full range of practice, and will affect panel relationships: Ford intends to deepen relationships with firms that demonstrate productivity gains and reduce its reliance on laggards. The unresolved question is pricing, and Croley is blunt about it: clients are still seeing annual rate increases without a persuasive account of how firms' AI investment produces any rate relief. Axiom's survey of 528 in-house legal leaders points the same way. Among AI users, all plan budget growth, yet 83% lack established ROI measures and only 7% are optimising and measuring AI across the department. Asked who they prefer for AI-enabled legal work, respondents chose ALSPs over law firms by 52% to 24%.
So what
The competitive unit is becoming the verified outcome, not access to a model. Clients are changing their provider mix before many traditional firms have changed their operating or pricing model, and the field they are choosing from now includes software vendors, law firms, ALSPs and supervised service businesses all competing across the same workflow. We cannot answer that by citing licence numbers or a handful of isolated pilots. It needs evidence that a defined matter type is genuinely faster, cheaper or better, plus a pricing mechanism that shares some of the gain with the client rather than banking all of it. The comforting half of the story is that the buyers have the same discipline problem we do: growing budgets with no task-level ROI leaves them unable to tell useful experimentation from expensive activity, which is exactly why the ALSP preference is a lead we can still close rather than a settled verdict. Law firms will need to first understand the impact of AI on matters and then transparently discuss this with clients - covering things like baseline time and cost, AI-enabled time and cost, the quality checks applied, the exceptions, and the pricing consequence for the client. One honest review of one matter type is worth more to a panel review than any amount of "we use AI extensively". New entrants such as Norm can burn through venture money to capture a market for only so long until they have to up their prices as well, law firms need to keep up or risk having sections of their work go to these aggressive new business models.

The Price of "Intelligence"

Artificial Lawyer - Harvey increases token use 14x in 6 months|Sourcery - interview with Harvey co-founder and head of infrastructure|OpenAI - Introducing GPT-5.6|Benedict Evans - Ways to think about token pricing

What
Harvey co-founder Gabe Pereyra said the platform's token consumption has increased 14-fold in six months, excluding embeddings. The underlying scale is roughly one trillion tokens in January and a June run-rate of about 13 trillion, and he gave indicative workload costs to go with it: a single assistant query can cost around $20, and reviewing 100,000 contracts can cost around $20,000. These are company examples rather than audited unit economics, and the $20,000 figure is a run across 100,000 contracts, not one contract review, but they make the production-cost question concrete for the first time. OpenAI's general release of GPT-5.6 makes the tiering just as concrete. Luna is priced at $1/$6, Terra at $2.50/$15 and Sol at $5/$30 per million input/output tokens. OpenAI reports that Sol improves quality on difficult professional work, and that programmatic tool calling can cut prompt-token use, supported by customer evaluations from Clio and Legora. Those are vendor and customer-reported evaluations, but on the DD benchmark GPT-5.6 Sol performed well, coming second (although under GPT-5.4 - and at a higher cost).
So what
The wrong response to any of this is a flat token ceiling. A cheap task can waste real money at scale, and an expensive task can still be excellent value, so a blanket cap punishes the wrong thing. The measure that actually matters is cost per verified outcome, with review, rework and failed runs counted in, not the raw token number on the invoice. Benedict Evans supplies the wider market caution worth carrying: supply, demand, marginal cost and buyer ROI are all still unstable, and the current dynamics point towards foundation models becoming commodity infrastructure with more of the value captured by the products and services built on top of them. That strengthens the case we have made before for model portability and routing, and against letting our long-term operating model depend on today's token prices or any one provider's frontier tier. Up until now, businesses encouraged staff to offload the work they disliked, but often failed to redesign the jobs or decide what the recovered capacity was actually for. So token governance is really a workflow and workforce question. We need bounded experimentation, task-level budgets, and routing rules that send routine work to cheaper models while escalating the difficult or high-risk work, and we need to decide what people will do with the time this gives back. A useful action would be to measure one workflow end to end, route it by difficulty and risk, and record its token cost alongside the human review and rework it generated, rather than reporting tokens as a standalone infrastructure line. First, think about what we can charge for this work or what the value is - then compare numbers. Firms with solid data architecture and captured information on what work costs to deliver today and how that is recovered and charged for will be at an advantage when thinking about pricing the value of output against token costs.

Worth Reading

Everything else worth a click

- Legal Market and Delivery

CMS introduces an AI value contribution bonus

Every employee below partner (bar the AI-build team) can earn up to £5,000 for demonstrable AI use. Introduced late 2025 and now public, it puts CMS alongside Simmons and Shoosmiths in paying staff to adopt. The open question is what "value contribution" is actually measured on, which is the ROI-proof problem DD2 and DD3 both land on.

Centari launches External Views

White-labelled deal-intelligence dashboards that turn an internal knowledge asset into a client-facing product. The same "productise your know-how" move as Cooley's GO Lab, one layer up.

DocuSign's legal-tech strategy

Intelligent Agreement Management now serves 40,000 companies and contributes 12% of ARR, and a proposed certificate of agentic action tackles provenance and auditability head-on. Worth reading for the provenance idea alone.

Legal IT Insider - Knowledge Exchange takeaways

Legal CIOs on data strategy, resilience, AI providers moving into services, and the continuing weight of human judgement. A good read of where firm-side infrastructure thinking has got to.

Spellbook launches Autonomous Contract Management

A late 30 June release, not covered in Edition 44, of a drafting product expanding into end-to-end contract workflow. Artificial Lawyer called it a "CLM killer"; take that with salt, but the direction is the point.

Headline - AI Europe 100

A business-model taxonomy that separates copilots, system builders and outcome engines, and locates Harvey and Legora among the system builders where the moat is embedded workflow and data. More useful than a sector map for thinking about where value sits. #wearethebeaver

Thomson Reuters and Anthropic on high-stakes professional AI

Promotional, but genuinely useful on the production bar for professional AI: citation self-checking, stability across long tool chains, and dependable context management. TR reports one internal remediation tool cut a root-cause investigation from three hours to four minutes.

Unhyped AI - what serious bank AI adoption looks like

An FS piece with a clean Access to Absorption to Control maturity model. Every bank figure in it is company-reported, so attribute rather than treat as independent, but the framing is the most workshop-ready diagnostic we hold right now.

Law Insider - How Transactional Lawyers Are Adopting AI in 2026

A 534-lawyer, 75-country survey of AI adopters. Usage is settled (86% use AI on contracts weekly) but trust is not (no tool clears one-in-five "very confident"), and verification and sources is the single most-requested feature. Produced by SimpleDocs, whose own tool tops most trust metrics on a community that skews solo and SME, so discount the tool league table and read the structural findings: the demand for grounding, and how far AI still trails on institutional-knowledge tasks.

- Policy, Courts and Governance

Demis Hassabis - A Framework for Frontier AI

Proposing a US-led, industry-funded frontier-AI standards body, benchmark-defined frontier models, voluntary pre-release testing moving towards mandatory approval, and post-release vulnerability work. The AGI framing is expansive and commercially interested; the institutional proposal underneath it is concrete enough to be worth the read. Axios summary.

FCA publishes the Mills Review

Billed as the first regulator-initiated review of its kind, on agentic AI in retail finance, consumer protection and seven recommendations for regulators and industry, with FCA research that around 11 million UK adults are likely to use agentic AI for personal finance. The strongest UK financial-services angle of the fortnight and directly relevant to FS clients.

OECD Employment Outlook 2026

The flagship annual, with around 28% of OECD jobs in occupations at high automation risk and a fresh graduate-unemployment hook. The authoritative labour read for the fortnight, and useful macro context for the token-and-jobs argument in DD3.

Civil Justice Council - update on AI and court documents

The consultation has closed, and the 30 June update points towards no new AI-specific requirements for professional legal drafting, proportionate transparency for expert evidence, and further work on witness statements and litigants in person. Useful context for the UKJT deep dive.

- Models, Agents and Work

Legal Benchmarks leaderboard

Independent, system-level evidence that usefulness and reliability are not the same metric: across 64 tasks the leader reaches 67.6% drafting reliability under an all-pass standard, and the most useful model manages 41.2%. The demanding, system-specific caveat is real, but this is the honest counterweight to any "polished output equals dependable output" claim.

OpenAI - GPT-5.6

General availability, three price-performance tiers (Luna, Terra, Sol) and programmatic tool calling, resolving Edition 44's gated restricted-preview item. The clean statement of what production-grade capability now costs per tier.

Meta releases Muse Spark 1.1

Meta's second Superintelligence Labs model (agentic, multimodal, 1M-token context) and its first paid API, at $1.25/$4.25 per million tokens and pitched on "aggressive pricing" to undercut OpenAI and Anthropic. It also landed well on legal work, topping Vals AI's Harvey Legal Agent Benchmark at launch (a lead since eroded as scores updated) and placing second on their credit-agreements task. A frontier lab arriving cheap and competent is the commoditisation pressure DD3 and Benedict Evans describe.

SpaceXAI releases Grok 4.5

An "Opus-class" model for coding, agents and knowledge work, trained alongside Cursor and priced at $2/$6 per million tokens (EU access lagging to mid-July). It currently tops Vals AI's Harvey Legal Agent Benchmark, ahead of Fable 5. With GPT-5.6 and Muse Spark, it makes three frontier releases in a single fortnight, all landing in the price band DD3 calls the emerging competitive floor.

Benedict Evans - Ways to think about token pricing

The strongest external analysis on the economics: today's prices reflect a supply crunch, the equilibrium is genuinely unknowable, and the dynamics favour commoditised model infrastructure with value captured higher up the stack.

Lenny's Newsletter - how tech workers are feeling in 2026

5,920 respondents; burnout up from 44.7% to 55.7%, while 82% say AI makes them better at the job. The fear is workload and quality, not replacement (51% worry about doing more for the same pay versus 22% about losing their job). Product-heavy sample, so read as sentiment not census.

The Geek in Review - Beyond the Model: The New Baseline

A fictionalised but sharp treatment of the rebound effect: the time AI saves quietly becomes the new minimum before anyone feels the benefit. Pair with Lenny's survey and Sam Harden; do not present it as empirical evidence.

Ian Cooper - Coding Agents: Driving in Gears

Argues that human oversight is a dial set by certainty: low-certainty work needs smaller steps and faster review, higher-certainty work permits larger batches. Coding-specific, but the model transfers cleanly to legal workflow design and beats generic "human in the loop" language.

Simon Willison - understand to participate

The case for keeping enough technical understanding to supervise and shape AI-enabled work rather than accumulating "cognitive debt". Relevant to how we frame AI-assisted build work.

- Education and Capability

University of Chicago Law School AI strategy

A concrete 2026-27 pilot that bans devices in nine core 1L courses and closed exams, teaches legal writing before layering in AI, requires an oral defence for substantial upper-level papers, and permits supervised clinical use. The useful formulation is learning without, with and about AI, and it is the sharpest attempt yet to preserve unaided reasoning while still teaching responsible use. LawSites analysis.

Ropes & Gray's AI competition for associates

Junior lawyers get billable-time credit for structured AI learning and reusable workflows. A neat way to make upskilling count against the metric juniors are actually measured on. Likely paywalled, so use only if accessible at drafting.