AI DEPRESSION SCORES DEPEND ON THE AI
Depression has no diagnostic blood test. Large language models promise tireless, inexpensive, repeatable ratings drawn from clinical conversations, and early evaluations have been encouraging. Psychiatry has faced a similar question before: in 1971, a study found that American and British psychiatrists disagreed after watching the same videotaped interviews, and the field answered with explicit criteria and structured interviews.
But a language model only becomes a “rater” through a series of choices. Someone picks the model, words the request — as a questionnaire, in a psychiatrist’s voice, or as the patient answering about themselves — chooses digits or letters for the answers, decides whether the model reads only the patient’s words or the whole conversation, and turns its output into a score. Evaluations rarely show these alternatives.
880 ways to build one rater
Baihan Lin, of the Icahn School of Medicine at Mount Sinai in New York, crossed 11 open models (from 1 to 21 billion parameters) with five wordings, two views of the conversation, two ways of reading the answer and four ways of building the score. That makes 880 raters, each applied to the same 189 recorded interviews, held by an animated virtual interviewer. Before each interview, participants had filled in the PHQ-8, an eight-item depression questionnaire scored from 0 to 24; a score of 10 or more is the usual screening threshold. About 30% of participants crossed it.
The study was pre-registered: the analysis code was frozen before any comparison with the questionnaire. The locked procedure was then run once on 86 new interviews led by a fully autonomous virtual interviewer, with automatic transcripts. All processing stayed on a local workstation.
The rater outweighs the patient
- One interview, opposite verdicts. A participant with mild symptoms (PHQ-8 of 5, below the threshold) was flagged as screening positive by half of the 880 raters. No participant got a unanimous verdict.
- Accurate raters still disagree. Among the 540 raters with good discrimination (an AUC of 0.70 or more), two drawn at random disagreed on the screening decision for 40% of participants.
- The model matters more than the person. The choice of model explained 30.0% of the variation in scores; stable differences between participants explained 10.5% — 2.8 times less.
A scale with a misplaced zero
Accuracy, as usually measured, says how well a rater ranks people. It does not say how many it flags. Raters flagged anywhere from 0% to 100% of participants, against 30% in reality. What set that share was each rater’s average tendency to over- or under-rate (R² = 0.91). The author compares it to a bathroom scale with a misplaced zero: it may rank people well, but its offset decides who crosses the line. Two models with nearly identical accuracy, Qwen2.5 7B and Mistral 7B, would flag 24 and 80 people out of 100, respectively.
Small technicalities moved that line. Reversing the answer letters — so that “A” meant “nearly every day” instead of “not at all”, a clinically meaningless change — shifted the share flagged by a median 21 points, and up to 85. Running the same model as a compressed file in different software changed a median 21% of decisions.
The ratings also picked up more than depression. They tracked post-traumatic stress symptoms as closely as depression, talkative participants were flagged more often, and 16% of raters reversed the direction of the small difference between women and men reported by the questionnaire.
Calibration helps, up to a point
Since the offset drove so much, the author tried a repair: resetting each rater’s threshold with 40 locally labelled participants. Accuracy rose from about 60% to 75%, and disagreement was halved, from 40% to 20%. But even when every rater was forced to pick the same number of people, they still picked different people about one time in five.
In the 86 new interviews, five of the six pre-registered predictions that could be tested held. The author lists clear limits: only open models up to 21 billion parameters, because the data could not leave the workstation; a self-reported questionnaire as criterion, not a diagnosis; no clinicians rating the same transcripts; and English-language interviews from a single laboratory.
Five checks before trusting a score
The paper ends with practical advice: report the range of decisions across sensible configurations rather than a single number; fix a neutral configuration in advance; calibrate locally and measure the disagreement that remains; test what else the rater picks up; and treat any change of model, software, transcription or wording as a new instrument. As the author puts it, when changing how we ask changes whom we identify, the instrument becomes part of the explanation.
