Medicine & healthPreprintExperiment3 min read

WHEN THE AI IS WRONG, A SECOND AI HELPS YOUNG DOCTORS

AI models can now look at a radiograph and the patient’s history and propose a diagnosis with a written rationale. Whether that helps depends not only on the model, but on how doctors react — especially when the suggestion is wrong. Earlier studies, cited by the paper, found that less experienced doctors gain the most from AI help but are also the most likely to be led astray by its errors.

A team from Nanchang University and the Hong Kong University of Science and Technology tested a simple idea drawn from collective intelligence: show doctors two independent AI opinions instead of one.

A randomized trial with three arms

Lin Wu, Zhe Xu, Fuqing Zhou and colleagues ran the trial at three hospitals in China between July and September 2026, registered with the Chinese Clinical Trial Registry. They recruited 132 residents with less than three years of experience, both in radiology and in other specialties (internal medicine, general surgery, emergency medicine), and randomly assigned them to three groups:

  • Group A: suggestions from GPT-5.4 alone;
  • Group B: GPT-5.4 plus Kimi-K2.6;
  • Group C: GPT-5.4 plus Gemini-3.6 Flash.

The participants did not know which models they were seeing. Nine residents dropped out because of network failures or feeling unwell, leaving 123 in the analysis.

All read the same 60 chest and abdominal radiographs. For each, they first gave a diagnosis on their own and locked it in, then saw the AI suggestion or suggestions, then gave a final answer. The cases were deliberately chosen so that each model was right on exactly 39 of 60 (65%) — and GPT-5.4, the shared model, was wrong on 21. For comparison, five radiologists with one to five years of experience scored 54.3% without AI.

What a wrong AI does

Across everyone, accuracy rose from 46.3% to 61.5% with AI help. The interesting part is what happened on the 21 radiographs where GPT-5.4 was wrong:

  • Radiology residents with one AI ended at 20.0% accuracy on those cases. Without AI, the same group had scored 31.4%: the wrong suggestion pulled them down.
  • With two AIs, radiology residents reached 40.1% and 40.4%.
  • Residents from other specialties show the same pattern, from 12.1% with one AI to about 31% with two.

Over all 60 cases, radiology residents improved by 9.6 percentage points with one AI, and by 16.3 and 17.5 points with two — about 7 points more. Reading time and confidence did not differ between the groups.

Not a free lunch

The second opinion has a cost. For residents outside radiology, when GPT-5.4 was right, adding a second model lowered accuracy from 87.4% to about 75%: a dissenting AI could talk them out of a correct answer. Overall, their improvement did not differ between one and two AIs. And when both models were wrong, an exploratory analysis found no significant benefit from the second one.

The authors conclude that two suggestions “may” limit the influence of erroneous AI, with a gain overall only for radiology residents.

Limits to keep in mind

  • The 60 cases were selected to give every model the same accuracy, so they do not reflect everyday practice.
  • The effect may depend on how often the two models disagree, which varies from one pair to another.
  • No formal sample-size calculation was made; the number of participants was based on feasible recruitment.
  • The random allocation itself was produced by asking another language model, DeepSeek, for a balanced list, which was used without modification.

Next, according to the authors, come studies with unselected patients and eye-tracking, to see how doctors actually weigh two disagreeing machines.

Legal notice