Search papers, labs, and topics across Lattice.
This study introduces a novel unbiased method to estimate the prevalence of LLM-assisted writing in biomedical publications by analyzing changes in word frequencies. The findings reveal that by the end of 2025, an estimated 89% of biomedical papers will exhibit LLM-associated vocabulary, with a notable disparity in usage between sections—68% in Discussions versus 32% in Methods. These insights highlight the growing influence of LLMs in academic writing and underscore the need for policy frameworks to address potential misconduct and fraud.
By 2025, nearly 90% of biomedical papers may be infused with LLM-generated language, raising urgent questions about academic integrity.
Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies.