AI-ERA PAPERS SAY "RATHER THAN" SEVEN TIMES MORE OFTEN
Large language models now help write a good share of scientific papers, and they leave traces: favourite words, a certain tone. One habit has started to irritate researchers in natural language processing (NLP): sentences built on “X rather than Y”, which define a claim partly by what it is not. According to the paper, reviewers of the latest round of NLP conference submissions complained about it, and conference organisers now discuss “AI-trope-ridden” writing.
Olga Zamaraeva, Adrián Gude, Roi Santos-Ríos and Carlos Gómez-Rodríguez, at the University of A Coruña in Spain, set out to measure it — with a title that makes the point by itself.
Seven times more
They compared 4,744 papers from the 2019 meetings of the Association for Computational Linguistics, before AI writing assistants were common, with 1,821 NLP papers posted on arXiv in 2026 in the same conference format.

Occurrences per 1,000 words in 2019 (blue) and 2026 (orange) papers: “rather than” and “X, not Y” surge, while “instead of” and “as opposed to” decline. — Figure 1, Zamaraeva et al. (2026), arXiv:2610.10092.
- “Rather than” rose from 0.134 to 0.993 uses per 1,000 words, about seven times more.
- The related “X, not Y” quadrupled; meanwhile “instead of” and “as opposed to” became less common.
- 94% of the 2026 papers contain “rather than”, against 36% in 2019, and they use it about eight times on average. One 2026 paper uses it 69 times, roughly once every 170 words; no 2019 paper exceeds 12.
The team also had models write complete papers from the title and abstract of real ones. GPT-5.6-sol and GPT-6-sol used the phrase in every paper, more often than the 2026 authors. Anthropic’s Claude Haiku 4.5, Sonnet 5.5 and Opus 5.5 used it in nearly every paper as well, at somewhat lower rates. An open model, Tülu 3 8B, rarely used it at any stage of its training. Revising papers did not remove the habit: later arXiv versions of the same 2026 papers contain slightly more.
Annoying, but which ones?
Frequency alone proves nothing: the phrase has plenty of legitimate uses. So two of the authors, experienced reviewers, labelled 1,000 uses shown blind, as “annoying” or “legitimate”.
- Almost no 2019 use annoyed them (1.6% for one, 0% for the other).
- About one 2026 use in ten did (11.5% and 8.3%).
- They rarely agreed on which uses were annoying. But since a 2026 paper contains so many, the authors estimate that about half of 2026 papers contain at least one use that annoys each of them, and nearly two-thirds one that annoys at least one of them — against about 1% in 2019.
What sets the annoying uses apart is mainly how they treat the rejected option. They tend to praise one alternative and belittle the other, and other people more often judged the rejected option a straw man, a position no competent researcher would defend. Readers also took the phrase itself as a sign of AI writing, while being unable to tell GPT-written passages from human ones written in 2019.
The Claude results, labelled so far by one annotator only, are mixed. Claude’s first drafts annoyed no more than 2019 papers and less than GPT-6-sol’s. Its revisions annoyed as often as GPT-6-sol’s, about twice as often as its own first drafts. The difference came from “disavowals”, in which authors decline a stronger claim about their own work (“a hypothesis rather than a tested mechanism”), which became more frequent when the models revised their papers.
Rewarded by the raters
Where does the habit come from? Chatbots are fine-tuned on pairs of answers, one of which a human or an AI judge prefers. In four public datasets of such preferences, answers containing “rather than” were more likely to be chosen in three, most clearly when the judges were humans. Four open “reward models” scored a sentence higher with its “rather than” clause than without it, even when a neutral clause of the same length replaced it. And telling a model to be “honest” made it use the phrase slightly more.
The authors conjecture that the construction is a side effect of this kind of training: in a single chatbot answer, a careful disavowal looks honest; repeated through a whole paper, and a whole literature, it reads as advocacy. Banning the phrase would probably just move the same function elsewhere, as GPT-6-sol does when it swaps “rather than” for “X, not Y” during revision.
The style is spreading
The study has clear limits: two annotators who are also authors, and the Claude comparison so far labelled by only one of them; pattern-matching that may miss uses; and preference data that only show correlations. But its conclusion goes beyond one phrase. Human-signed papers have taken up the construction, so the style of the scientific literature is drifting towards that of the models — which will, in turn, learn from it.
Conflict of interest. The authors state that they used Claude Code (Claude Sonnet 5, Sonnet 5.5 and Opus 5.5) to write their code, run the analyses and draft text throughout the paper, including the Methods and Results sections, before editing it by hand — removing, among other things, ten “rather than” introduced during generation. Claude models are also among the systems studied and served as automated raters. Claude also wrote the present article.
