Search papers, labs, and topics across Lattice.
HKUST(GZ
3
0
6
Current scientific agents struggle to maintain a coherent narrative across evidence and calculations, with only 34.81% achieving strict accuracy in complex tasks.
Static benchmarks mislead model performance assessments, as LiveHouse-TS reveals dramatic shifts in rankings when evaluated in real-time environments.
LLM multi-agent systems can achieve significantly higher accuracy at a fraction of the cost by learning to selectively delegate tasks instead of relying on rigid orchestration.