Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
PRO-LONG boosts LLM performance on long-horizon tasks by 18% while using up to 5.8 times fewer tokens than traditional methods.
Over a quarter of tasks in popular AI benchmarks contain critical flaws that distort model evaluations, and this automated auditing framework can catch them.