Issue #217

Why Legal Adopted AI Agents First, Not Engineering

Legal's Codex usage jumped 108x since February, far outpacing engineering's 5x—and Google's first industry AI targets finance and law.

BusinessWhy Legal Adopted AI Agents First, Not Engineering

Last Thursday’s issue told you how the legal AI company Harvey built its own model in just two months. There was one question I left unanswered then. Why legal, of all fields?

Half the answer had already surfaced three days earlier. On August 25th, Google Cloud unveiled its first industry-specific Gemini Enterprise offerings — and there were only two. Finance and legal.

The other half is buried in enterprise usage data OpenAI updated on August 12th. It’s a table showing how many times weekly active Codex users among enterprise customers multiplied by job function, from February 2026 through June. The top spot wasn’t engineering.

It was legal. 108x. Engineering came in at 5x.

Codex is a coding tool. Lawyers don’t write code. So that’s exactly where I want to start today — from this mismatch. This isn’t a story about legal suddenly warming up to AI. It’s a story about where agents actually take root first: not in the field that loves AI most, but in the field where “what counts as a good result” is already written down on paper.

Before You Take the 108x at Face Value

Let me first look at this number properly. It’s indexed to February 1, 2026 as 1, plotted through June.

FunctionMultiple vs. February
Legal108x
Sales41x
Recruiting41x
Marketing26x
Healthcare24x
Finance & Accounting20x
Engineering5x

Even a16z, who put out this chart, added its own caveat: “how much of this shift owes to Codex’s own broader rollout is worth scrutinizing.” I want to add one more line to that. This is a multiple, not an absolute figure.

legalIf Codex users in the legal function were effectively at zero in February, then a 108x multiple could come from just a small increase in absolute numbers. Conversely, engineering’s mere 5x isn’t because its growth was slow — it’s because its starting point was already high. If you ranked these by absolute user counts, the order would almost certainly flip. Reading this as “legal overtook developers” is a misread.

So what actually survives from this table? I’d argue relative ranking does.

Codex’s broader rollout wasn’t opened to legal alone. It opened equally to sales, marketing, and finance. If everyone started from the same floor and legal still hit 108x while marketing only reached 26x, that 4x gap can’t be explained by “the door opened.” It’s a question of who walked through first once the door opened. Today’s piece is my attempt to answer exactly that one question.

One more thing worth noting. In the same report, as of June, Codex accounted for 64% of combined Codex-plus-ChatGPT output tokens among enterprise customers. That means more tokens went toward handing off tasks to be completed than toward conversations seeking answers to questions. What the legal function switched on wasn’t a coding tool — it was an agent.

Google’s first industry-specific versions are finance and law

On August 25th, Google Cloud released Gemini Enterprise for Financial Services and Gemini Enterprise for Legal simultaneously, in preview. They’re the first two industry-specific versions. Healthcare and life sciences are listed only as “coming soon.”

Both products share the same structure: four layers — industry-specific domain skills, MCP1 connectors, agents that actually perform the work, and a partner ecosystem.

Finance’s flagship offering is the Financial Research agent. It ships with more than 50 built-in skills and handles tasks like credit risk assessment, portfolio monitoring, KYC, and bond issuance. Google says it’s cut the time to analyze risk exposure across a complex bond portfolio to under 5 minutes. The data it connects to comes from sources like FactSet, Moody’s, MSCI, PitchBook, Guidepoint, Dun & Bradstreet, and SEC EDGAR.

On the legal side, it covers contract review and redlining, playbook2 drafting, regulatory monitoring, legal research, and DSAR3 response. It connects to iManage, NetDocuments, DocuSign, Everlaw, RelativityOne, Thomson Reuters HighQ, and even the free case-law database CourtListener. Development involved Cleary Gottlieb, Freshfields, Weil, and Williams & Connolly, while Deutsche Bank and CME Group were brought in on the finance side.

Up to this point, it’s an unremarkable enterprise product announcement. What I lingered on was the next sentence.

When the finance agent produces a result, it delivers a confidence score, an explicit methodology, an audit-ready data snapshot, and precise source citations alongside it. And in the data governance section, there’s this line: “Licensed data remains licensed. Permissioned data remains permissioned.”

The key point is that this isn’t a newly written product spec — it’s professional norms transcribed verbatim.

Agents Arrived First in Jobs Where “Passing” Is Already Written Down

Overlay two signals and the picture snaps into focus.

Law and finance aren’t fields known for loving AI. If anything, they’ve long had a reputation for being conservative. And yet these two fields have three things no other field has.

First, the output is text. Contracts, memos, opinions, research notes — exactly the form an agent can produce.

Second, what counts as “passing” is already written down. Law firms have clause-by-clause playbooks; banks have KYC checklists and underwriting criteria tables. Decades of regulations and precedent have piled up. That’s precisely the material you need to train and grade an agent. Last issue I mentioned that Harvey built 1,750 task environments and 75,000 grading items by hand to train its model. Law is the field that already had the master copy of that scorecard from the start.

Third, showing your reasoning is already a professional norm. Lawyers never argued a point without citing sources to begin with. Analysts always disclosed their methodology. So “cite your sources, state your methodology explicitly, keep an audit trail” isn’t a safety feature bolted onto an AI product after the fact — it’s the output spec this profession already demanded.

legalaiPut these three together and here’s what you get. In marketing, “is this good copy?” is hard to reach consensus on. But in law, “does this clause violate our playbook?” is a verdict you can render. When there’s a passing standard, you can verify whether the result is correct — and once you can verify it, you can hand the task to an agent.

So when Google picked finance and law as its first targets, it wasn’t just about market size. It was choosing the industries where an agent’s output can be graded. In industries where grading isn’t possible, nobody can tell whether the agent did well or badly.

The most interesting item on this launch’s list of legal connectors is who’s on it. Harvey and Legora are both there.

The two companies are the headline rivals in the legal AI market. Harvey was valued at $11 billion this past March; Legora at $5.6 billion in April. Yet instead of pushing them aside, Google folded them in as connectors. In Legora’s case, once both sides’ customers sign off, some of Legora’s features become usable inside Gemini Enterprise, with Legora’s source citations and grounding4 left intact. The trade publication Artificial Lawyer read this as “not competition but ecosystem.”

I see it a little differently. The moment a company becomes a connector, the screen users open every day is Google’s product, and the legal AI company becomes the party supplying features from behind that screen. Rather than shoving competitors out, the platform absorbs them as partners and takes over the point of contact with users.

And Harvey has already answered this. As I covered in a previous issue, Harvey shipped its own model, Tenet, just two months later, staking its mission on letting “law firms own their own intelligence.” The timing makes the point unmistakable: Harvey announced on August 18, Google on August 25. A week before its name showed up on Google’s connector list, Harvey had already declared it would not become a component.

So right now, competition in legal and financial AI is being fought across three layers: models, evaluation environments, and data access.

LayerWho holds itWhat’s being sold
ModelGoogle, OpenAI, AnthropicIntelligence itself. A commodity remade every couple of months
Evaluation environmentHarvey, LegoraThe ability to know what counts as a passing answer
AccessFactSet, Moody’s, iManage, Thomson ReutersThe raw data and the permissions attached to it

What Google has shipped is a product that stitches these three layers together under a single permissions system. Last issue, I framed it as “the model is the commodity, the environment is the capital asset.” This announcement adds one more line to that: after the environment, access is the next most valuable thing. No matter how good an agent is, if it can’t read documents inside iManage under the correct permissions, a law firm simply won’t use it.

Korea already has 3 tracks running

There’s a reason this doesn’t sound like a distant story.

Law firm Bae, Kim & Lee (Taepyeongyang) became the first Korean law firm to adopt Harvey company-wide in July, then in August became the first major Korean law firm to roll out ChatGPT Enterprise company-wide. Domestic case law runs separately, through LBox and Superlawyer. International contracts go to Harvey, domestic precedent goes to homegrown legal tech, and document drafting and translation go to general-purpose AI. That’s already 3 tracks.

What Google is selling sits exactly on top of those 3 tracks — a governance layer. Not a fourth tool to add to the pile, but a layer that makes the growing pile of tools operate under the same permission rules.

So the question Korean organizations need to ask right now isn’t “which AI should we buy.” It’s two questions.

First: does your organization’s repetitive work have a scorecard?

A contract-review playbook, review criteria, a checklist. Without one, you can’t judge whether any tool’s output is actually good. If you can’t judge it, you can’t delegate it, and a human ends up re-reading everything anyway. That’s why law and finance moved first. Building this scorecard isn’t an AI project — it’s a job-definition project, and it costs working hours from the business side, not IT budget.

Second: does the system you’d attach an agent to actually hold permissions?

If you’ve ever worked at an organization where internal documents are scattered across personal PCs and messaging apps, you’ll immediately understand why the line in Google’s announcement — “permissioned data stays permissioned” — is the core of the product. The connector to a law firm’s system works because that firm uses iManage. If a document has no permissions attached to it, there’s simply nowhere to attach anything. For most Korean organizations, the bottleneck isn’t model performance — it’s this.

Closing

Let me sum it up in three points.

First, that 108x figure for Codex users in the legal profession is heavily distorted by a base-rate effect. Still, the relative ranking — 4x ahead of the next-closest profession under the same conditions — is a real signal.

Second, Google’s choice of finance and law as its first industry-specific targets comes down to this: these two industries already have a standard for grading an agent’s output.

Third, the competitive battleground has shifted from model performance to data access rights and permission systems. The proof is that Google folded Harvey and Legora in as connectors — and that Harvey, just a week earlier, had answered with a model of its own.

Here’s one thing I’d suggest trying this week. Pick the three tasks that repeat most often in your organization, and for each one, write five lines describing what output would count as a “pass.” If you have those five lines, you can verify whether a result passes, and that means you can actually hand off the task. Without them, all you’re doing is piling up tools. Law got there first because it already had those five lines written down.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

  • OpenAI, “Enterprise signals: What frontier firms are doing differently”, updated 2026.8.12. ··· This is the primary source for the 108x figure. Worth reading alongside the per-role multipliers is OpenAI’s own definition of “frontier firms” (the top 10% of monthly AI usage) and its caveat that tokens are an imperfect proxy for business value.
  • a16z, “Charts of the Week: Winds of Thematic Change”. ··· The post that introduced this chart. What matters is that the authors themselves flagged the question of “how much of this is just the Codex rollout expanding.”
  • Google Cloud, “Introducing Gemini Enterprise for Legal”, 2026.8.25. ··· Check the connector list yourself — Harvey and Legora are both on it. Chapter 4 of today’s issue came directly from that list.
  • Google Cloud, “Introducing Gemini Enterprise for Financial Services”, 2026.8.25. ··· This is where you’ll find the 50+ built-in skills, the confidence scores and audit snapshots, and the line “licensed data stays licensed.”
  • Artificial Lawyer, “Google Launches Gemini Enterprise for Legal”, 2026.8.25. ··· Evidence that the industry read this announcement as “ecosystem,” not “competition” — the opposite of my own reading, so it’s worth weighing both side by side.

Background


Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. MCP (Model Context Protocol): A standardized protocol for how AI connects to data and functions in external systems. Rather than wiring up a custom connection for every service, it’s closer to using an outlet that fits a standard plug.

  2. Playbook: An internal reference table that a law firm or corporate legal team uses to pre-define, clause by clause, what’s acceptable and what must be revised. For a human, it’s a manual; for an AI, it becomes a scoring rubric.

  3. DSAR (Data Subject Access Request): An individual’s right to ask a company, “tell me what personal data you hold on me.” Because there’s a response deadline and it requires combing through an entire organization’s documents, it’s a classic example of work that’s expensive to do by hand.

  4. Grounding: A method that ties an AI’s answers to actual source documents — and requires it to cite them — so it doesn’t fabricate responses. If no supporting document can be found, the safer choice is for the AI not to answer at all.