Search papers, labs, and topics across Lattice.
6
0
8
A new framework categorizes molecular LLM agents into four levels of autonomy, revealing critical gaps and risks in their deployment for scientific discovery.
Achieving similar performance to larger models with significantly less data and faster inference speeds could redefine efficiency benchmarks in foundation models.
Hallucinations in chemical reasoning models coexist with correct answers, revealing a complex relationship that challenges our understanding of model reliability.
Adversarial purification can be dramatically improved by focusing on patch-level semantics, leading to state-of-the-art performance in defending against adversarial attacks.
A single structural edit can drastically impair LLM performance in molecular tasks, highlighting the fragility of their generalization capabilities.
PhySciBench reveals that top LLMs struggle with scientific reasoning, achieving only 33.5% accuracy, while DelveAgent demonstrates a promising 7.5% improvement in performance.