AI can give us another answer in seconds. That does not tell us whether it deserves to change a decision.
By the time a legal representative in Victoria read a report about a young child, it had already travelled some distance.
A child-protection worker had prepared it. Their team manager had reviewed and signed it. It had been submitted to the Children’s Court.
About a week later, on the day of another hearing, the legal representative noticed that something did not fit. The language was unusual. The assessment of the risks around the child was inadequate. They suspected ChatGPT had been used and raised the concern.
The investigation that followed found a disturbing contradiction. Elsewhere in the report, a doll had been described in connection with sexual behaviour by the child’s father. Later, the same doll appeared among the parents’ strengths: evidence that they were supporting the child’s development with “age-appropriate toys”.
The department withdrew the report and replaced it. The child’s eventual care and court decisions did not change, so this is not a story about an AI error changing a child’s life.
But it was evidence being placed in front of people who could.
The obvious response is that somebody should have checked it.
Somebody had.
Another answer does not guarantee a better one
The comforting response to unreliable AI is to keep a human involved.
The evidence is less comforting.
Researchers brought together more than 100 experiments comparing people working alone, AI working alone, and people working with AI. On average, the combination performed worse than whichever of the human or AI was better on its own, with the losses concentrated particularly in decision tasks.
That does not mean AI is generally better than people. It means that putting two sources of intelligence together does not automatically give you the best of both.
Sometimes the mistake also runs in the opposite direction.
In Aberdeen, Yvonne Cook went for routine breast screening and took part in an NHS Grampian evaluation of an AI system. The normal double-reading process did not recall her. The AI flagged something.
It did not diagnose her or make the final decision. Its disagreement triggered another human review.
Yvonne was called back. A small Grade 2 tumour was confirmed and treated. Across 10,889 women in the prospective study, the same additional-review process detected 11 cancers that routine double reading had not recalled.
The important part is not that a machine beat two doctors. Two readers had reached one conclusion; another form of intelligence produced a reason to look again, and a third person reconsidered the case.
In Victoria, a human challenged a representation influenced by AI. In Aberdeen, AI challenged an existing human judgement.
The question cannot simply be who gets the final word.
What makes disagreement worth listening to?
There is a clue in research on judges making bail decisions in the United States.
Most judges who departed from an algorithmic recommendation did worse than they would have done by following it. A small minority did better. Those better-performing judges were more likely to use relevant information about the particular case that the model did not possess, and less likely to be pulled around by something merely vivid or salient.
That distinction travels well beyond courts.
If another intelligence disagrees with you, what does it know that you do not? If you disagree with it, what do you know that it does not? And does that difference matter enough to change the judgement?
Confidence is not new evidence. Seniority is not new evidence. Neither is a fluent answer from a machine.
West Midlands Police offers the darker version. Before the decision to exclude Maccabi Tel Aviv supporters from a match at Aston Villa, some material gathered through AI search had already entered a developing case. Later, while the justification was being strengthened under challenge, a written briefing cited a match between Maccabi and West Ham that had never happened. Official reviews found confirmation bias and inadequate challenge; a conduct investigation remains live.
A judgement does not only respond to evidence. It can start shaping which evidence gets looked for.
There is a profound difference between a judgement that tests itself against the evidence and one that recruits evidence to support itself.
When the answer starts shaping the record
This becomes more important when AI is not simply producing an answer at the end of a process. It is beginning to participate in creating the evidence people will later judge from.
Think about an ordinary clinical consultation. A patient speaks. A clinician listens. Afterwards, the encounter becomes a record.
That compression is not new. Handwritten notes never contained everything that happened in the room either.
Ambient voice technology changes who participates in deciding what survives.
If something false is added, somebody may still see it and remove it. If something material is left out, the next person may act as though it was never said.
What is left out can shape the next judgement as much as what is written down.
The room matters too. A straightforward follow-up with one patient and one known problem is different from a conversation involving a relative, an interpreter, uncertainty, mental health or safeguarding.
Same tool, different room.
This is happening while the technology can offer genuine gains. England’s independent patient-safety investigator, HSSIB, has nevertheless already opened an investigation into ambient voice technology in hospitals. It says adoption is accelerating while the safety implications are not fully understood, and that research and implementation have concentrated more heavily on efficiency than patient-safety risk. Its report is expected in summer 2027.
The benefits can become visible to measurement before some of the harms become visible to the system.
That is not an argument to stop. It is a reason to keep the judgement open as the evidence develops.
The person inside the record
There is another observer we tend to leave out of the human-versus-machine argument.
The person the record is about.
Years before generative AI entered consulting rooms, OpenNotes began giving patients in the United States access to what clinicians had written about them. In one large study, around one in five people who had read their notes reported something they believed was wrong.
That figure alone needs caution. Believing something is wrong does not prove that it is.
But in a smaller study where patients could report potential problems, clinicians judged many of those concerns to be definite or possible safety issues. In more than half of the cases subsequently confirmed with patients, the record or the care changed.
The patient is not the final judge of the medicine. But they may know whether they actually take the medication listed in front of them, which side hurts, what they said in the room, or what happened after they left it.
The information sets are different.
Patients can already see substantial parts of their GP record and, increasingly, some hospital documents through the NHS App, but secondary-care information remains incomplete and varies between providers. The direction is towards people being able to see more of the evidence through which their care represents them — and, in some services, increasingly contribute back into it.
Often, today, there is no such observer.
Victoria had one because another professional happened to read the report and ask a different question. Aberdeen did something more deliberate: disagreement itself triggered another reading. OpenNotes created a route by which the person represented could sometimes put missing information back into the case.
Good judgement cannot depend on the right person happening to arrive at the right moment.
What the answer deserves to change
This is where Crump’s Law becomes a judgement problem.
Once an encounter has passed, the record may become what later people judge from. The same can happen with a report, a risk score, a photograph, an old assessment or a reputation. Technology makes those representations easier to retrieve, combine and reproduce, while the original situation moves on.
Getting another intelligent answer is becoming easier too.
What remains difficult is deciding what that answer deserves to change.
Return to Victoria. The important thing about the legal representative was not that a human defeated an AI; another human had already reviewed and signed the document. Aberdeen does not support the opposite conclusion either.
What mattered was that the judgement had not yet become completely closed to something material that did not fit.
Victoria had a lawyer with a different question. Aberdeen had a process that forced another look when two forms of intelligence disagreed. OpenNotes gave some patients a way to put information back into the record before later decisions were made.
The harder future may not be one in which we lack intelligent answers.
It may be one in which answers arrive so easily that we stop asking what would deserve to overturn them.
What would actually change your mind?
Sources
Office of the Victorian Information Commissioner — Investigation into the use of ChatGPT by a Child Protection worker (2024)
https://ovic.vic.gov.au/wp-content/uploads/2024/11/DFFH-ChatGPT-investigation-report-20240924-Re-upload.pdf
de Vries, Lip, Staff et al. — Prospective evaluation of artificial intelligence integration into breast cancer screening in multiple workflow settings: the GEMINI study, Nature Cancer (2026)
https://www.nature.com/articles/s43018-026-01126-1
Vaccaro, Almaatouq & Malone — When combinations of humans and AI are useful: A systematic review and meta-analysis, Nature Human Behaviour (2024)
https://www.nature.com/articles/s41562-024-02024-1
Angelova, Dobbie & Yang — Algorithmic Recommendations and Human Discretion, NBER / Review of Economic Studies
https://www.nber.org/papers/w31747
House of Commons Home Affairs Committee — The policing of the Aston Villa v Maccabi Tel Aviv match (2026)
https://publications.parliament.uk/pa/cm5901/cmselect/cmhaff/1553/report.html
Bell et al. — Frequency and Types of Patient-Reported Errors in Electronic Health Record Ambulatory Care Notes, JAMA Network Open (2020)
https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2766834
Bell et al. — A patient feedback reporting tool for OpenNotes: implications for patient-clinician safety and quality partnerships, BMJ Quality & Safety
https://qualitysafety.bmj.com/content/26/4/312
Health Services Safety Investigations Body — The use of Ambient Voice Technology in hospitals (2026)
https://www.hssib.org.uk/patient-safety-investigations/the-use-of-ambient-voice-technology-in-hospitals/
NHS App help — Viewing documents
https://www.nhs.uk/nhs-app/help/documents/
NHS England Digital — Hospital and specialist appointments in the NHS App
https://digital.nhs.uk/services/nhs-app/nhs-app-features/hospital-referrals-and-appointments-in-the-nhs-app
NHS England — Single Patient Record – your health at your fingertips
https://www.england.nhs.uk/digitaltechnology/the-single-patient-record/
