Search papers, labs, and topics across Lattice.
Affiliation:
4
0
6
LLMs rarely ignore context simply because it contradicts their parametric memory, meaning reported context-memory conflicts may largely be artifacts of poor LLM-as-a-judge selection.
VetScore reveals that risk-weighted evaluations can significantly enhance the reliability of veterinary QA systems, achieving expert-level alignment even with minimal model complexity.
Coding agents can outperform raw data models in time series analysis, but still miss 22-34% of questions, revealing critical reasoning gaps.
Small LLMs paired with symbolic solvers can outperform larger zero-shot LLMs on formal reasoning tasks, but still struggle with multilingual inputs.