Search papers, labs, and topics across Lattice.
This study evaluates the efficacy of advanced browser-based large language models (LLMs) in extracting nuanced data from scientific literature through a series of four workflows. The findings reveal that while LLMs can generate effective prompts autonomously, they still struggle with context interpretation and autonomous literature discovery, often leading to missed or hallucinated references. Ultimately, the research outlines a collaborative framework where experts set standards, models perform cross-checks, and human oversight resolves discrepancies, facilitating scalable scientific data curation without sacrificing quality.
LLMs can autonomously generate prompts that rival expert-written ones, but still fall short in accurately interpreting scientific context and discovering literature.
Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the performance of frontier, browser-based large language models (LLMs) to extract highly contextualized information. We demonstrate four escalating workflows, 1) given an expert curated prompt and research articles, most frontier LLMs perform well at data extraction, however can struggle with interpreting scientific context and nuance, 2) given simple instructions, LLMs can author their own prompts which were almost as eNective as expert-written prompts, 3) autonomous discovery of research literature was diNicult, agents either missed or hallucinated references, and 4) LLMs can create new datasets from published guidelines that closely match human-expert judges, but still require a human-in-the-loop. Together, these findings define an auditable division of labour in which experts specify the evidence standard, models cross-check repeated extractions and researchers resolve disputed cases, providing a practical route to scaling scientific data curation without relinquishing expert oversight.