Skip to content
← Back to Insights

The Asymmetry of Trust: Why We Forgive Human Error and Fear the Machine's

Human error scatters. Model error repeats. That is not a difference in how bad each mistake is, it is a difference in shape, and it is most of what looks like a double standard.

August 12, 2026 · Quantum Nexus Ventures FZCO

A judge has a bad afternoon. Tired, distracted, carrying an unconscious bias they would deny if asked, they read a case file a little less carefully than the one before it. The ruling is wrong, or at least worse than it should have been. Nobody writes about it. There is no headline, no regulatory inquiry, no public reckoning about whether judges as a category can be trusted with the job.

An AI system assisting that same court hallucinates a citation once. It becomes a story. Sometimes it becomes a law review article. Sometimes it becomes the reason an entire pilot programme is paused.

The instinct is to call this unfair, and it is worth taking that instinct seriously rather than dismissing it. But the honest answer is not that we are being irrational about AI. It is that two failure modes that look similar from a distance are not the same shape at all.

The part that is bias

Some of the asymmetry really is just how people process risk, and it has been studied since the late 1970s. The psychometric paradigm, set out by Fischhoff, Slovic, Lichtenstein and colleagues in 1978 and summarised in Slovic's 1987 review in Science, found that lay judgments of risk track two dimensions rather than expected fatalities. Slovic labelled them dread risk, defined by perceived lack of control, dread, catastrophic potential, fatal consequences and an inequitable distribution of risks and benefits, and unknown risk, defined by hazards judged unobservable, unknown, new, and delayed in their harm. Dread is the dominant of the two. The finding is not that people cannot count: asked directly, laypeople estimate annual death tolls roughly as experts do. Those estimates simply are not what their sense of risk is made of.Sources: Slovic, Perception of Risk, Science 1987 · Fischhoff et al., 1978

Per passenger mile, driving is around a hundred times deadlier than scheduled commercial aviation: 7.28 against 0.07 fatalities per billion passenger miles in Ian Savage's United States figures. Roughly 39,000 Americans die on the roads each year. The gap narrows sharply on other denominators, and measured per journey rather than per mile the two modes are much closer, which is worth saying because the choice of denominator is doing a lot of the work. Yet one airliner crash draws investigative and regulatory attention that no equivalent count of road deaths attracts.Sources: Savage, fatality risks across United States transportation modes

A hallucinated citation lands on both dimensions at once. It is new and poorly understood, which is unknown risk. It arrives through a system somebody else chose, and it concentrates into a single nameable event rather than a diffuse pattern spread across a career, which is dread. Judged by that psychology alone, the reaction is predictable and not really about AI. It would happen to any new, opaque, involuntarily encountered risk.

That is real, and it explains part of what looks like a double standard. It does not explain all of it.

The part that is not bias

A biased or exhausted judge produces an idiosyncratic error. It happens in this case, on this afternoon, shaped by this specific person's specific bad day. It is bounded. It does not automatically recur in the next case. Over a career those errors scatter, some in one direction and some in the other, and the system absorbs the noise imperfectly through appeal, through precedent, through the slow attrition of professional reputation.

A model with the same underlying flaw does not scatter. It repeats. The same misreading of the same kind of clause, the same blind spot in the same category of question, fires identically every time the pattern recurs, across every matter the system touches, simultaneously, until somebody notices. The failure is not one bad afternoon. It is every afternoon at once, silently, for as long as it takes to detect.

That is not a claim about which is more accurate. On many measures the machine is better on average than the tired judge. It is a difference in the shape the error takes when it does occur, and you cannot see that shape by looking at any single instance in isolation.

None of which is our observation. Reliability engineering has called it common-cause failure for decades: dependent failures that defeat the very redundancy introduced to improve reliability, and which are estimated to account for a large share of safety-system unavailability in nuclear plants. Kleinberg and Raghavan formalised the algorithmic version in 2021, showing that decision-makers converging on one shared algorithm can collectively do worse than decision-makers using idiosyncratic, individually less accurate processes, with no external shock required to produce that result. Bommasani and colleagues showed the corollary a year later: when systems share components, the same people get refused everywhere.Sources: Kleinberg & Raghavan, Algorithmic monoculture and social welfare, PNAS 2021 · Bommasani et al., outcome homogenization, 2022

One boundary on that argument, because it matters to anyone holding a single file. Uncorrelated error averages out across many independent decisions. It does not average out for the person it lands on, and no amount of aggregate scatter helps the client whose case was the bad afternoon. The distinction is about what the system as a whole can absorb, not about what any one matter can.

The concern is not hypothetical, though it is easier to describe than to count. Citation fabrication in litigation is now documented in the open: a public database of court decisions identifying it recorded over 1,600 cases by mid-2026, of which around 650 involved lawyers rather than self-represented litigants, and its compiler is explicit that the count is a floor and skewed towards jurisdictions whose filings are easy to search. The portfolio-level analogue, a compliance system missing the same category of violation across an entire book of business because the defect sits upstream in a retrieval or classification step with no named owner, has no published case study we are aware of. That absence is itself the point. A failure mode with no owner is also a failure mode with nobody positioned to report it.Sources: AI Hallucination Cases database

The apparatus we stopped noticing

There is a second reason the comparison feels unfair, and it has nothing to do with psychology. We did not arrive at our tolerance for human error by deciding to be lenient. An enormous amount of infrastructure was built to catch, correct and price it, over centuries, until that infrastructure became invisible through sheer familiarity.

Appellate review exists because trial judges get things wrong. Malpractice insurance exists because professionals make mistakes that cause real harm. Licensing boards, peer review, publicly reported disciplinary records, the entire apparatus of professional accountability, all of it was built specifically because human judgment is fallible and somebody decided that fallibility needed a system rather than an apology.

None of that reads as scrutiny anymore. It reads as background. It is old enough, and distributed widely enough across institutions already trusted, that nobody experiences it as anyone being watched. It is still there, still doing the work, every time a case is appealed or a licence reviewed.

AI systems do not have that apparatus yet. What looks like an unreasonable demand for perfection is, underneath, a demand for the same kind of infrastructure human institutions already have: a record of what happened and why, a way to trace a bad outcome back to its cause, a named party who answers for it, and a mechanism for correction that does not depend on somebody happening to notice. The demand is not that AI never be wrong. It is that it be wrong inside a system built to catch it, which is quietly what is asked of every human professional whose work carries consequences.

Where this actually leads

The standard will not disappear once that infrastructure exists. It will stop feeling like a standard, the way appellate review stopped feeling like an accusation against judges. The scrutiny AI systems face now is not evidence that the bar is unfair. It is evidence that the accounting is still being built in public, one gap at a time, in exactly the place where the equivalent human infrastructure has had a few hundred years to fade into the furniture.

The question worth asking is not whether AI will ever be forgiven the way a tired judge is forgiven. It is whether the record it leaves behind can do, from the first day, what centuries of appeals and licensing boards eventually learned to do for human judgment: turn an individual failure into a traceable one, and a traceable one into a correctable one, before it repeats itself silently across everything the system touches next.

This is an opinion / thought-leadership piece. It is not legal or financial advice.