Same Facts, Different Answer: Legal AI’s Consistency Problem

By Shaz Aziz, Neota Logic.

It is August 2026. A lawyer with no technical background can now build a working legal AI tool in a chat window over a coffee break. I have done it myself. It is revolutionary, and anyone in legal tech who pretends otherwise is selling something.

So, I will happily concede the point that many vendors try to argue around: the models are good. They are excellent, even. They read contracts, they extract obligations, they summarise, they converse. For the parts of legal work that never fitted neatly into rules, they are the best tools we have ever had.

The question that should keep lawyers awake is not whether AI can produce the answer. It is what happens in month six. The model has been updated twice. The person who wrote the prompt has left. A regulator, or opposing counsel, asks how a particular decision was made, and whether the same facts would produce the same decision today. In most legal AI deployments I see, the honest answer is that nobody knows.

The generator and the validator

There is a persistent and seductive story that each new wave of technological progress destroys the last. The iPod killed the CD player. Then Apple themselves destroyed the iPod with the iPhone. The music players kept changing over the years, but the music simply moved from home to home. The technology that mattered was never the device. It was the thing the device carried – structured recordings that any new player can ingest.

The same misreading is happening in legal AI: logic systems are the old thing, language models are the new thing, and progress means throwing away the old one in favour of the new one. It confuses the music player with the music collection. But the rules are the music. A logic system holds codified judgement that organisations have spent years making explicit: playbook thresholds, escalation rules, compliance logic. The chat window that the judgement comes through is the current device, and LLMs are another change of home. Buyers who feed years of codified judgement into a prompt are heading into deployments that will come back to bite them.

Neota Logic.

There are two kinds of machine here and they are good at different jobs. A language model is probabilistic. It produces a plausible answer, usually a very good one, and a slightly different one each time. A rules engine is deterministic. Same inputs, same answer, every time, with a path through the logic you can print out and hand to someone to inspect. One machine reads and converses. The other makes decisions and remembers why. Legal work needs both, usually in the same workflow: the model reads the document, the rules make the decision, and where the risk warrants it the workflow stops for a person and will not move on without them.

Every recent “hallucination” case that fell foul of professional standards had a human in the loop somewhere: a paralegal, a clerk, local counsel. Presence is not a control. A control says what the person must verify, evidences that they verified it, and blocks the next step until they have.

Where do your rules live?

When a client asks me why they cannot simply build their compliance logic inside their AI assistant (and they do ask), I have stopped arguing about capability. Instead I ask one question: where do your rules live?

If the answer is “in a prompt”, consider what that means. A prompt can sit in version control, but that versions the words, not the behaviour: the model underneath changes without asking you, so the same prompt in January and in June is not the same system. A rules engine versions the behaviour itself. There is no maintenance process: rules generated by prompting get maintained by reprompting, and a reprompt can quietly rewrite the whole thing. The auditor’s question is clear and simple. Can you reproduce a decision made in January under the rules as they stood in January? For a prompt-built system the answer is no, and no amount of model improvement fixes it.

Neota Logic.

A large insurer or firm is not making one decision, it is making tens of thousands a day. Make those systematically, with inspectable logic, and you have an operation. Make them probabilistically, with logic that can shift under you unannounced, and you are making inconsistent decisions at scale. That is the pattern class actions are built on: when one decision is successfully challenged, every decision made the same way is in play, and you cannot show they were made the same way. One of the claims applications running on our platform has roughly 109 million possible paths through its decision logic. Every one of them is inspectable. That is not a property you can prompt your way to.

What comes next

The future I would bet on, and the one we are building, has three stages.

First, AI working inside governed workflows. The model does the reading, extraction and classification; the rules make the decision; the audit record shows which steps were probabilistic and which were deterministic. This is shipping today and it is where serious deployments start.

Neota Logic.

Second, the direction of travel reverses. Instead of AI inside the workflow, the workflow becomes callable from wherever the lawyer already works, such as a chat interface. Ask a question with consequences, and rather than the model guessing, it invokes a deterministic application and returns an answer with the reasoning attached. The model makes exactly one decision: to call the workflow. Everything that has to be repeatable stays on the other side of that line. The rules engine stops being a place you go and becomes something your assistant carries, the way your phone carries your music.

Third, AI helps author the rules themselves, which makes the deterministic layer cheaper to build and maintain. That is a direction rather than a product today, and I would be wary of anyone who tells you otherwise.

The question to ask

Almost nobody in legal AI can tell a buyer which part of an answer was a guess. That is the question I would put to any vendor, us included: point at the workflow and say which steps are probabilistic, and which are rules. The extraction might still be wrong (it is a model, and models are wrong sometimes), which is exactly why the workflow stops there for a person. But you know where the boundary is. That is worth more than any amount of reassuring language about safety and trust.

So when you evaluate legal AI this year, skip the model benchmarks. Ask where the rules live, who can change them, and whether they can reproduce last January’s decision under last January’s logic. The vendors who can answer that will be happy you asked.

Learn how Neota Logic keeps your legal AI answers consistent here.

About the author: Shaz Aziz is Senior Director, Client Solutions, at Neota Logic.

[ This is a sponsored thought leadership article by Neota Logic for Artificial Lawyer. ]


Discover more from Artificial Lawyer

Subscribe to get the latest posts sent to your email.