Search papers, labs, and topics across Lattice.
The paper introduces STRIVE, an LLM-based framework designed to automate the generation and evaluation of controlled event sets for psycholinguistic studies, specifically targeting event plausibility judgments. By varying a single event slot while keeping other features constant, STRIVE enhances the efficiency of creating these sets, achieving a significant improvement in quality from 16.7% to 75.0% with the addition of a global reasoning scratchpad and evaluator-guided refinement. Despite these advancements, the study reveals that events near the plausibility boundary remain challenging, highlighting the necessity for human input in certain conditions.
STRIVE automates the generation of event plausibility sets, achieving a remarkable 75% quality rate, but still struggles with boundary cases that require human judgment.
Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausibility effects, these studies require controlled event sets in which one event slot varies across plausibility levels while all other event features remain fixed. Constructing such sets manually is labor-intensive. We therefore introduce STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets crossing plausibility class (plausible vs. implausible) with intended classification difficulty (easy vs. hard). Given a verb, STRIVE constructs a shared event frame, then produces one event per condition by varying one slot while holding all others fixed. In experiments with six models across 60 verbs, GPT-5.1 produced high-quality sets only 16.7% of the time using the baseline generation prompt. Adding a global reasoning scratchpad and evaluator-guided refinement raised this rate to 75.0%. Greater reasoning effort also improved evaluator--human agreement. Nevertheless, events near the plausibility boundary remain most difficult. They elicit the greatest human disagreement, and the best evaluator reaches only 57% accuracy on the implausible-hard condition, indicating a need for human input. Overall, STRIVE offers a scalable approach to reducing manual effort by automating initial event-set generation and evaluation for psycholinguistic studies.