Search papers, labs, and topics across Lattice.
Fudan University
2
0
3
Even state-of-the-art language models struggle significantly in real-world tasks, exposing critical shortcomings in their deployment readiness.
GPT-5's scientific reasoning skills plummet by nearly 50% when tackling multi-step workflows, revealing a critical gap in current LLM agents' ability to orchestrate complex tool use.