Search papers, labs, and topics across Lattice.
This paper introduces PhysicsBench, a comprehensive benchmark and leaderboard designed to evaluate generative and predictive AI models in engineering design and simulation using a standardized procedure across various tasks and data scales. By assessing 66 models on nine datasets, including industrial-scale simulations, PhysicsBench provides a more realistic evaluation framework that emphasizes geometric fidelity and physical accuracy under constrained data conditions. The findings reveal that traditional academic performance does not reliably predict success in practical applications, as the top-performing models vary significantly with data scale and task type.
Traditional academic benchmarks fail to predict real-world performance, as PhysicsBench reveals that the top models shift dramatically across different data scales and tasks.
Generative and predictive artificial intelligence models are increasingly used to generate geometry and to predict physical fields and scalar quantities in engineering design and simulation. Yet these models are typically evaluated in isolation, on academic datasets at unconstrained scales, with inconsistent metrics and procedures. We present PhysicsBench, a unified benchmark and leaderboard that evaluates generative and predictive models under one standardized procedure. PhysicsBench spans seven generation and prediction tasks across 1D, 2D, and 3D domains and ranks 66 models on nine datasets, comprising industrial-scale CAD/CFD/FEA simulations and public references, expanded into 28 configurations. One procedure and ranking apply to both families, each ranked within its own tasks. Evaluation spans realistic, limited data scales from S to XL rather than the unlimited training sets common in academic benchmarks. A common metric suite captures geometric fidelity with distributional distances, physical-field and scalar accuracy, and engineering-specific field- and shape-validity. BenchRank debiases correlated metrics and ranks by PageRank over a head-to-head dominance graph, so every reported quality metric is also ranked, with computational cost in a separate efficiency view. Across tasks, an architecture's large-scale academic standing weakly predicts its small-data ranking. The top model changes with data scale in six of the seven tasks, and no model leads more than one task. PhysicsBench turns"state-of-the-art"from a self-reported claim into an openly published foundation for model selection.