Search papers, labs, and topics across Lattice.
This study evaluates six frontier language models (LLMs) against a new benchmark, EuroExec, consisting of 413 open-ended European executive decision tasks, using over 4,000 hours of human expert evaluation. The findings reveal that the best-performing model only achieves a 56.9% solve rate, significantly lagging behind human experts who solve tasks at near-ceiling levels and are preferred in 74% of direct comparisons. This highlights a critical gap between LLM capabilities and expert judgment in complex decision-making scenarios, emphasizing the need for human evaluators in assessing model performance on nuanced tasks.
Frontier LLMs struggle to match human expert performance, solving less than 57% of complex European executive tasks.
Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on. We dedicate more than 4,000 human expert hours to evaluate a selection of six frontier LLMs on a member of this class of problems: EuroExec, our introduced human expert-based benchmark composed of 413 open-ended long-form European executive tasks authored by 47 vetted domain experts, each question drawn from experience in a real case. Every response is manually evaluated through a multi-attribute rubric, an item-specific checklist of requirements, and a preference rank ordering, extracting an aggregate metric"Solve Rate". The strongest model solves only 56.9% of tasks, while expert-written reference answers judged blindly are solved at near-ceiling levels and are preferred over every model response in 74% of direct rankings, placing frontier generative systems well below the professional standard of work they are already used for. We see that the best way to extract this kind of conclusion is by employing human evaluators, carefully checking their consistency through rigorous statistical analysis, and observe that automatic measurements also fall short when evaluating on this case of real-world open-ended problems with a subjective ground truth.