Search papers, labs, and topics across Lattice.
2
0
5
0
This work compares open-weight and API-served general-purpose, multimodal, and astronomy-specialized models using judge-based correctness and complementary reference metrics and highlights the role of domain-specific evaluation in determining which models, capabilities, and evaluation criteria are appropriate for scientific workflows.
Achieve up to 4.79x higher throughput in LLM serving by dynamically switching between data and tensor parallelism on the fly, without restarting workers.