top of page
Professional Woman Smiling

AI: Brutalize Your Own Argument

PL_Architecture_LinkedIn.png

Why Every Lawyer (and Every Judge) Should Be Running Their Reasoning Through a Trained AI Workflow Before It Costs 

Partners and Associates — and, Oh My, Judges — Hallucinate Too

​​The standard objection to AI in law is that it hallucinates. This is true. So do junior associates at 11 p.m. on a Sunday, so do senior partners who skim the headnote and bluff the rest, and so do paralegals who have been told the deadline is crushing on them. The profession has lived with confident-but-wrong for as long as the profession has existed. Lawyers have just historically billed for it.

​

The standard response to AI hallucination, after the Sixth Circuit’s sanctions opinions and the Ninth Circuit’s mandamus in Lnu v. Blanche, has been to ban the tools, to use them timidly, or to use them in secret and pretend the brief wrote itself at 2 a.m. the way briefs have always written themselves at 2 a.m.

​

All three responses miss the point. The point is not whether AI can write a brief. The point is whether AI can stress-test a brief—your brief—before opposition does, before the trial judge does, and before the Court of Appeal turns the premise you skipped into a published reversal with your name in the caption.

​

It can. The tooling exists. The workflow is reproducible. The only live question is whether the profession adopts it deliberately and openly, or keeps muddling through with consumer chatbots until the next sanctions opinion is named after one of us.

Understand layers of architecture- its simple

The architecture has three layers, and lawyers who get this wrong tend to mix them up the way clients mix up the IRS and the FBI. 

​

​The first layer is the foundation model itself—the underlying large language model that does the reasoning. Today that means Anthropic’s Claude, OpenAI’s GPT family, Google DeepMind’s Gemini, or an open-weight model you host yourself (Llama, Qwen, Mistral, take your pick) when you do not want your work product crossing a vendor’s servers. 

​

The second layer is the agentic platform or wrapper that sits on top of the model and gives it tools, memory, instructions, and the ability to act in sequence—Claude Code, OpenAI’s Codex, Cursor, Perplexity’s research interface, your firm’s in-house build, or any of a dozen others that exist as of this morning and may not exist by Christmas. 

​

The third layer is the legal research environment the platform is allowed to read: Westlaw, Lexis, Bloomberg Law, Fastcase, vLex, CourtListener, the firm’s knowledge, the docket pulls, the deposition transcripts, the regulatory record. 

The architecture matters more than the AI brand

Mixing the three layers up—pretending Perplexity is a model or that Westlaw is an agent—is the tell that the speaker has read about this technology but has not used it. The architecture matters more than the brand on any invoice.

Your Own Litigation Stack — Your To-Do List

You give the platform custom instructions, a skill library, and subagents — each subagent with a job description and no feelings to hurt. You build the stack by talking to it, out loud or in text.

​

First, you train the bench. You unleash your agents to study — and you tell each one who to be:

​

  • One studies civil (or criminal or administrative) procedure and every seminal case, and is instructed to think like a trial judge ruling in real time and like an appellate panel reviewing that ruling on a cold record.

  • One studies every case pertinent to your area of expertise — for us, consumer protection, toxic torts, developmental toxicology, expert admissibility — and is instructed to think like an appellate justice writing the opinion the panel will actually publish.

  • One studies the doctrinal architecture above your field — due process, separation of powers, preemption, standing, the administrative state, the First Amendment overlays — and is instructed to think like a Supreme Court justice scanning for the structural argument nobody below noticed.

  • One studies evidence and scientific methodology — Daubert, Sargon, Sanchez, Kelly/Frye where it still bites, — and is instructed to think like the gatekeeping judge who will decide whether your expert testifies at all.

  • One studies ethics, sanctions, and the professional-responsibility rules — and is instructed to think like the disciplinary panel reviewing opposing counsel's conduct, and yours.

​​

Then, you spin up the case-specific stack:

​

  • One parses the statutes and rules that govern the motion in front of you.

  • One reads the cases and pulls the actual holdings, sorted ruthlessly from the dicta the losing side will quote at you.

  • One verifies every citation and quote against the primary source — the agent that exists because nobody else on the team wants the job.

  • One thinks like the appellate panel that will read the record you are building: preservation, standard of review, the gap between the record you have and the record you wish you had.

  • One thinks like a Supreme Court justice scanning for the structural argument the panel will miss and the policy hook the dissent will build a career on.

  • One is instructed to be brutal to your position — every weakness, every unstated assumption, every authority you are stretching past its holding like a sweater on the wrong dog, every fact that cuts against you and that you were quietly hoping nobody would notice.

  • A separate hawk-eye agent audits the first stack and asks the question nobody on the trained team will: what did you make up, and what did you skip because it was inconvenient?

​

Then, the discipline: you run the same brief through a second stack entirely. Different foundation model, different agentic platform, ideally a different vendor across both layers—pair Anthropic’s Claude through one wrapper with OpenAI’s Codex through another, or any commercial model against an in-house open-weight build, whatever combination forces the two passes to actually disagree instead of nodding at each other like in-laws at Thanksgiving. Same instructions. Same brutal-critique mandate. You ask them to triangulate, too. Where they diverge, you have located the soft tissue of your argument—before opposing counsel does it for you, in a footnote, in front of the panel:

​​

  • Rinse and repeat in another stack, running on different models.

  • Have each stack audit the other's work.

  • Sit down and run your own audit — read every case, every citation to the record.

  • Expect more questions than you started with, and send the stacks back in, again and again.

  • Only when you are fully satisfied, run one last audit through the legal AIs you (marginally) trust. Tell agents to highlight all edits. Check again.​

​

​And then—only then—you do the work no AI does and no AI should. You verify every source the system relied on. You read the cases yourself, not the summaries of the cases. You confirm the quotes at the pinpoint, not at the headnote. You apply law to your facts, your client, your jurisdiction, in your own voice—the voice you will use when the panel asks the question you were not expecting. You ask agents to build a table of every authority with summaries and context for each quote, which you now verify manually- this is the same table that will sit on the lectern at oral argument and save you when the presiding judge interrupts with “counsel, in X v Y we said….” Only then do you sign the brief.

​

This is the line every court that has looked at the question has drawn, and it is worth saying in one sentence so nobody pretends it is more complicated than it is.

​

AI assistance is permitted. Do not sign anything you have not personally read against the raw source. The verification standard is not citator-green. It is “I pulled the case, I read the cited passage, and I am satisfied that the authority supports the proposition for which I am citing it.”

​

If a junior associate, a contract attorney, a Brief Writer, or a chatbot drafted the paragraph, that is fine. The signature line is a representation that you, personally, did the verification. There is no theory of vicarious AI liability. There is only signature liability.

This Does Not Replace You. It Frees You To Do More

The lawyers who fear this technology have it exactly backwards, and usually by midnight on a Sunday. The grunt work—the fifth re-reading of a string cite, the hunt for the controlling-but-not-quite-on-point case, the eighteenth pass through the record looking for the preserved objection that absolutely must be in there somewhere, please, God—that is the work AI does well, cheerfully, and at three in the morning.

 

The judgment work—which authority to lead with, which weakness to disclose and which to bury politely, which expert to retain, which witness to depose first, when to settle and when to take the case to verdict and ask the jury a hard question—that work remains entirely yours. It always will. Anyone who tells you otherwise is selling something or has never sat second chair.​

What changes is the floor.

The floor for a competent legal argument used to be a senior associate with two weeks, a Westlaw subscription, and the standing offer of pizza. The floor is now a senior partner with two days, the same Westlaw, and a stack of agents that have already run every citation through three verifiers and stress-tested the theory against the panel they think the case will draw. Your client gets the senior-partner brief on the senior-associate timeline at the small-firm rate. You get to actually think about the case instead of cite-checking your own prose at 1 a.m. and wondering whether “supra” is still a word.

The Colleague Who Actually Knows Your Case

Every lawyer was taught the same thing in the same words: bounce your thinking off a more experienced colleague. Say it out loud, watch them frown, listen for the question you did not want to be asked. Moot courts businesses are built on that premise. So is the hallway visit to senior counsel. So is the practice argument with a partner who has been doing this since B.C. and will tell you, without smiling, that your second point is your worst one.

​

It is good advice. It is also, in the cases that actually matter, structurally unreliable. The cases that matter are the ones built on your deepest, narrowest expertise—the ones where the doctrine is so specific, the regulatory record so long, the science so granular, that the field of qualified bounce-boards shrinks to a number you can count on one hand. And those few colleagues, however brilliant, are not living and breathing your case. They have not read the eighth deposition. They have not watched the defendant’s expert hedge on cross. They have not spent six months inside the regulatory file. By definition, they cannot see all the dark corners of your room. They are looking at it from the doorway, in low light, with their coat still on.

​

And there is the other problem nobody likes to say at the CLE dinner: sometimes the experienced colleague is just wrong. Confident, senior, well-credentialed, and wrong. The cases where your expertise is deepest are the cases where the bounce-board is most likely to give you outdated intuition wrapped in a confident tone. You walk out of the meeting reassured and worse-prepared, which is the most expensive combination in litigation.

​

A trained AI workflow does not replace that colleague. It does something the colleague structurally cannot. It reads every deposition in your case. It reads every order in your case. It reads the regulatory file at the granularity you read it, because you fed it the regulatory file. Then it brings to that record the full surrounding field—every published Prop 65 opinion, every related discovery sanctions case, every analogous safe-harbor ruling in any jurisdiction that has touched the issue. It does not light up one foggy room with a single flashlight held by someone who has been in the building maybe twice. It lights up the entire neighborhood, with your room at the center of the map, and shows you the structural beams behind your own walls.

​

That is the move. Not a smarter colleague—a colleague who has actually read your case, every page of it, and who can also read every adjacent case in the field in the time it takes to refill the coffee. The bounce-board problem was never that bouncing was wrong. It was that the surface you were bouncing against was too small, too unfamiliar with your record, and too occasionally mistaken to catch you when it mattered. The surface just got bigger, brighter, and a great deal better-briefed.

The New-Era Bench Drafts Opinion Before the Clerk Finds the Coffee Machine

Here is the part that will get me in trouble at every CLE dinner this year. Imagine if courts of appeal were running the same stack on incoming briefs.

​

Imagine a Ninth Circuit panel with an AI workflow that, on day one, flags every authority either brief cites that has been reversed, depublished, distinguished, or quietly limited in an unpublished disposition the lawyer pretended not to find—before the law clerk has opened a Word document or asked where the coffee is.

 

Imagine an appellate or trial court whose pre-argument memo includes a brutal-critique pass on both sides' strongest authorities, written by an agent that does not care which party prevails, does not need a weekend, and has no opinion about football. Imagine a trial court whose tentative ruling runs, before issuance, through a stress test that asks the questions a careful clerk would ask if a careful clerk had eight extra hours and were immune to the tall lawyer in the Hermès tie: does this ruling actually make sense, or was the bench charmed by the handsome dud with the good pocket square? Does it swallow one side's facially erroneous reading of the regs whole? Does it pretend the regulatory record stopped moving sometime around the first season of Friends?

​

A great deal of wasteful litigation, a great deal of remanded error, and a great deal of avoidable appellate work would simply not occur. Attorneys would see their flawed premises before committing them to PDF. Trial courts would catch the structural problems in their tentatives before the parties caught them in a writ. Appellate panels would identify the issues that actually needed briefing and could spend oral argument on questions that drove the outcome instead of questions the panel happened to find interesting at 9:07 a.m. The work product would be sharper. The judicial output would be more consistent. The cost to the client—and to the public footing the bill for the published reversal—would fall.

​

This is not a fantasy. The components exist today, run on consumer-grade hardware, and cost less than a single deposition transcript. They are imperfect, they require expert human supervision, and they will continue to require it for as long as anyone reading this still practices law. But they exist. The only question is whether the legal system adopts them deliberately—with disclosure rules, audit trails, and competence standards—or whether each generation of lawyer and judge re-learns the lesson Lnu v. Blanche just delivered to the Ninth Circuit bar, in slip-opinion form, with a citation reserved for future use.

What “Brutal” Actually Looks Like

The brutal-critique agent is instructed to identify, in descending order of how much it will hurt at the hearing:

​

Every authority cited that does not actually hold what the brief asserts it holds. Every quote that has been trimmed in a way that changes its meaning. Every statute the brief implies is dispositive but where the regulatory text would let a careful judge go the other way. Every factual proposition that depends on an expert opinion the brief does not yet have. Every standard of review premise that, if the trial court denies the motion, will lock in a deferential review on appeal. Every authority that is binding in the user’s home jurisdiction but persuasive elsewhere—and every place the brief assumes binding force without saying so.

​

The brutal-critique agent does not soften any of this. It does not say “you might consider whether perhaps...”
​

It says: “your reliance on Wilshire Westwood at p. 744 is structurally weak because that case holds on a contamination-not-exposure ground that opposing counsel will use to distinguish you into a string cite of your own. Move it to a string cite yourself, lead with the regulatory text, and put the burden on the defendant to explain why a 2026 record requires the trial court to ignore 30 years of toxicology.” It tells you what your friend would tell you if your friend had read every case in the database and owed you nothing. That is what “brutal” means.

​

Then the hawk-eye agent audits whether the brutal-critique agent missed anything—because brutal does not mean infallible, and overconfident AI is still AI. The hawk-eye hunts for: confabulated reversal history; misstated holdings; cases that look on-point but were procedurally limited in a footnote nobody quotes; opinions the primary agent cheerfully treated as published when they were depublished by the Supreme Court three months later; quotes verified at the wrong pinpoint by an agent that was, briefly, just guessing. The hawk-eye is the agent that catches the agent. Internal affairs, with no union.

​

Then a second stack runs the same brief, with the same instructions, on a different foundation model through a different agentic platform—the entire point being that the second pass has never spoken to the first and does not share its blind spots. You read both outputs side by side. Convergence is your reinforced ground. Divergence is your soft tissue, and where you should be most awake. You do the human verification of every authority you intend to cite, in the actual reporter, in the actual decision. Then you write the brief. Then you sign it. In that order, on purpose.

The lawyer who outsources judgment to the stack will only produce embarrassment faster

Three honest limitations, because anything that sounds too good in a CLE always is. First, none of this works without competent expert human supervision. The lawyer who outsources judgment to the stack will produce embarrassing work product faster, more efficiently, and at greater scale than they ever could before. The verification step is non-negotiable. The lawyers who skip it are the lawyers whose names show up in the next round of published sanctions opinions, and the published sanctions opinions are getting more specific every quarter.

​

Second, the tools are not yet equally accessible. A solo practitioner has different access to Westlaw, to the top-tier foundation models, to a clean agentic platform sitting on top of those models, and to the engineering time required to wire all three layers together than a litigation firm with an in-house technology budget and a director of innovation whose entire job is buying these things. That asymmetry will close over the next five years—the prices keep falling, the open-weight models keep gaining—but it is real today, and it is doing real harm today. The bar should think about it now, before it becomes a competence-floor problem under California Rule 1.1, and before “did you run it through a brutal-critique stack?” becomes the new “did you Shepardize?” in a malpractice deposition.

​

Third, judicial adoption is a separate problem from advocacy adoption, and it has its own ethics issues. A judge who uses an AI workflow to stress-test a tentative ruling is not delegating decision-making to the machine—any more than a judge reading a treatise is delegating to the treatise author, or a judge reading a clerk’s bench memo is delegating to the clerk. But disclosure, audit trail, who built the workflow, who tuned it, who reviewed its output, and who answers when it gets something wrong—those are real questions, and pretending they are not will not make them less real. Courts and judicial conduct commissions will have to work through them. The answer is not to ban the tools, because the tools are not going anywhere. The answer is to govern their use openly, with the same gravity we apply to every other thing a judge consults before issuing a ruling.

Why Not Just Buy Harvey?

The obvious question, somewhere around the third CLE on this topic, is why any of this needs to be built. There is, by the most generous count, more than a hundred legal-AI products on the market. Harvey raised at an eleven-billion-dollar valuation in March 2026 and now sits inside Paul Weiss, A&O Shearman, PwC, and KPMG, plus another 1,300 firms and legal departments who line up to be quoted in the next funding round.

 

Thomson Reuters bundles CoCounsel into every Westlaw seat that will hold one. LexisNexis sells Protégé the same way. Spellbook, Paxton, Legora, Vincent AI, Hebbia, and a long tail of two-engineer startups all promise the same thing in slightly different fonts: a friendly chat box on top of a foundation model, pointed at curated legal data, with a UI designed by someone who once read about a deposition.

​

I have tried them. I keep coming back to my own stack. This is not stubbornness and it is not the founder energy of someone who likes building things for the sake of building them. The honest reason is structural. A productized legal-AI wrapper is built for the median user at the median firm doing the median task. It is engineered, by definition, to be safe, defensible, and good enough for a deal-room first pass at a hundred contracts none of which carry your bar number into a courtroom. That is a real market and a useful product. It is not the market I practice in.

​

Build the bespoke stack if you want intensity at depth

​

Toxic cases at the level we run it is a deep-expertise field where the cases that matter turn on a single subdivision of a single regulation, a forty-year-deep active regulatory record, and a scientific question no published opinion has actually answered. The brutal-critique pass I need is not a generic “find weaknesses in this brief” button. It is a pass tuned to my own threshold for what counts as a weakness, my own taste for which authorities are worth distinguishing in text versus burying in a footnote, and my own habit of stress-testing the agency’s reading against a defense theory the defense bar has not yet articulated but will.

 

That threshold gets retuned, sometimes several times a day, when a deposition produces a new fact, when an order lands that I did not expect, or when a colleague reads a draft and frowns at the third paragraph. A productized wrapper cannot tune that quickly because it is not supposed to. Its job is to be consistent across thousands of users. My job is to be uncompromising for one.

​

There is a quieter reason as well. The big legal-AI products are built to be acceptable to general counsel and risk committees, which means they are built to refuse, to hedge, and to surround any sharp answer with a soft pillow of qualifications. That is appropriate for the deal-room use case. It is the wrong instinct for a contested matter where the value the system has to provide is exactly the sentence the soft pillow is trained to suppress—the sentence that says, in writing, that the argument the lawyer is leaning on is the weakest argument in the brief and the panel will eat it. A self-tuned stack will say that sentence. Harvey is not going to say that sentence to a Paul Weiss employee. They have different jobs.

​

None of this means the legal-AI industry is unserious. It means the industry is solving a different problem—the problem of bringing modern tooling to firms that are not going to staff their own AI workflow and would not know what to retune even if they did. That is a worthwhile problem, and the firms that buy those products will be better off than the firms that ban the tools.

​

For a litigator with a zeal and ability to better most creadential oppoonents  a strong opinion about their own threshold for rigor, and a willingness to spend all non-sleeping time wiring foundation models to a research environment that already holds every Statement of Reasons she cares about, the productized layer is too slow and one layer too many. The wrapper that ships in a box does not know your case. You do.

​

That is the trade. Buy the productized stack if you want competence at scale. Know you will be slower than many. Build the bespoke stack if you want intensity at depth. 

Lawyers who use AI to flatter their arguments will lose to lawyers who use AI to brutalize them

AI will not replace lawyers. Lawyers who use AI will replace lawyers who do not. That observation is now old enough to vote, and tired enough to skip the CLE keynote.

​

The new observation is sharper, and less comfortable. Lawyers who use AI to flatter their own arguments will lose to lawyers who use AI to brutalize them first. Judges who use AI to confirm their tentatives will be reversed by judges who use AI to interrogate them. The competitive advantage is no longer access to the tool—the tool costs roughly what a junior associate spends on coffee.

​

The competitive advantage is the willingness to point the tool at your own reasoning and sit with the answer, even when the answer is that your favorite argument was always weak and you knew it.

​

Be brutal to your own position first. Triple-verify your sources. Run the results through a second system that does not share the first one’s prejudices. Read the cases yourself, in the actual reporter, with the actual margin notes. Then sign the brief—because there is no theory of vicarious AI liability, and there never will be. There is only signature liability. Your bar number is on the line, not the model’s.

​

It is, in the end, what good lawyering has always been. Skeptical, source-verified, allergic to its own first instinct. The tools just got faster, and a great deal less tolerant than your junior associate.

​

bottom of page