Search papers, labs, and topics across Lattice.
This study evaluates the effectiveness of large language models (LLMs) in generating project plans for scientific research in physics, astrophysics, and cosmology, comparing outputs from human researchers and three LLMs. The evaluation involved blind assessments by human reviewers and newer LLMs, revealing that while human reviewers rated AI-generated proposals similarly to human ones, AI reviewers favored AI outputs. The findings highlight the potential of LLMs to assist in project planning but also raise concerns about biases in AI evaluation of proposals.
AI-generated project proposals can match human quality, but AI reviewers show a bias towards favoring their own outputs.
We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.