STILL FUNDED, NO LONGER COUNTED
Research funders increasingly learn what they fund by letting software classify the text of their grants. The US National Institutes of Health (NIH) have done so since 2008. Their RCDC system mines each award’s title, abstract, public health relevance statement and specific aims, and every year the agency reports to Congress and the public how much it spent in more than 300 categories, from cancer to minority health. Researchers, advocates and NIH staff all use these totals as the record of what science the public pays for.
Studies of such classifications assume that the applicant writes the text. Fangfang Xie (Nanjing University), Jingwen Zhang (Hubei University) and Haining Wang (Indiana University School of Medicine) asked what happens when the funder itself edits the text its software counts.
What changed in 2025
According to the documents the paper cites, executive orders of 20 and 21 January 2025 targeted diversity, equity and inclusion (DEI) programmes and “gender ideology”. On 4 March 2025, NIH guidance told program officials that awards which “do not support DEI activities, but may contain language related to DEI” should be updated “with the DEI language removed” before funds were released. A federal court vacated these directives on 16 June 2025; on 21 August the Supreme Court left that decision in place while staying the part that set aside grant terminations. In December 2025, NIH introduced an automatic text screen, reported to check at least 235 terms.
Following the words
The authors traced the names of population groups through 37,790 continuing awards, using public NIH files, and compared them with records the 2025 review did not touch: whom linked clinical trials enrolled, how linked articles had been indexed by human librarians, and whom trial eligibility criteria named. They focused on the category Racial and Ethnic Minority Health Research, and had some projects judged by reviewers blind to their identity and category.

The review acts on the summary, the count is mined from the summary; flagged continuations losing every screened term rose with the delay in issuing them. — Figure 1, Xie, Zhang & Wang (2026), arXiv:2610.03443.
What the numbers show
- The edits were massive. Of 5,461 continuing grants with a screened term in 2024, 39.7% had none left in 2025, against 0.08% the year before. The later a renewal was issued, the more it lost — up to 54.6% after more than 90 days. The edits read meaning: “a minority of cells” was rarely touched.
- The names were not decoration. In new grants with trials reporting enrolment, a title naming Black populations predicted a 60-percentage-point higher share of Black participants.
- The count followed the names. Among continuing projects carrying the minority health category, 99.4% kept it when the summary was unchanged, but only 15.7% when it no longer named a racial or ethnic population — while the projects kept their funding.
- The lost projects met the definition. Blinded reviewers judged that 42 of 50 randomly drawn losses had met the category’s definition in 2024. Comparing before-and-after summaries of 70 losses, they found the stated aims about the population rewritten in 49; read alone, 63 of the new summaries no longer met the definition. “The category read the edited text correctly.”
- A hole in the totals. NIH’s published total for the category fell from $2,612 million (2024) to $1,944 million (2025). In the authors’ analysis, continuing projects that lost the label after losing the name account for about a fifth of the net fall they could trace — a part invisible in the published figures.
Words, or people?
For sexual and gender minorities, the pattern was different. Of 2024 awards naming them, 46.7% were recorded as terminated at some point, against 4.8% of all awards — 37.8 points more after adjusting for vocabulary, institute and grant type, compared with +3.2 for Hispanic and +0.2 for Black populations. By March 2026, most of these were listed as possibly reinstated. Because listed and unlisted forms of the names were removed alike, the authors read this as population targeting: the pressure was on the group, not the words.
A record that cannot tell
The authors call the main result “uncounted science”: research the funder still pays for, that other records show involves the population, but that its own count no longer records. The problem, they stress, holds whatever the direction of the policy: when the words a funder counts are also the words it polices, its policy reads back as a fact about science.
Their limits are stated plainly: some tests rest on small numbers (20 strictly documented projects), reviewer agreement was sometimes only fair, edit dates and authors are unknown, and the window is too short to know whether the research itself changed. The authors declare no funding for the study. “A classification that reads text,” they conclude, “is only as valid as the governance of that text.”
