Search papers, labs, and topics across Lattice.
4
5
6
13
Backtesting LLMs reveals that recency can mimic leakage, leading to misleading scores that obscure true model performance.
LLMs leak future knowledge into past predictions, but a new method called TimeSPEC can filter this temporal contamination by verifying claims and regenerating predictions based on available information.
LLMs learn better when you evolve their system prompts alongside their weights, boosting generalization in reasoning tasks.
LLMs in gastroenterology can be made significantly safer: a new framework achieves near-human expert alignment and boosts accuracy by 8% via rejection sampling.