Search papers, labs, and topics across Lattice.
1
0
3
3
This work compares open-weight and API-served general-purpose, multimodal, and astronomy-specialized models using judge-based correctness and complementary reference metrics and highlights the role of domain-specific evaluation in determining which models, capabilities, and evaluation criteria are appropriate for scientific workflows.