Search papers, labs, and topics across Lattice.
Xinjiang University
9
0
10
Graph-augmented LLMs fail to leverage provided graph evidence effectively, revealing a critical gap in their design that can be addressed with targeted interventions.
ATLAS outperforms existing systems in medication safety for older adults, achieving a 53.73-point lead in decision-making accuracy over top proprietary models.
Current instruction-based video editing models are far from satisfactory, revealing critical gaps in evaluation that could reshape the field.
Pressure integration in humanoid motion imitation significantly enhances accuracy and stability, revealing the limitations of traditional vision-based methods.
DR-DCI achieves a remarkable 73.3% accuracy in agentic search tasks while efficiently scaling from 100K to 10M documents, outperforming traditional methods.
Continual learning for LLM agents hits a wall: scaling models doesn't reliably improve skill generation, and self-feedback can lead to recursive drift.
A clever two-stage agent using smaller models can produce better, more substantive peer reviews than brute-force application of the largest LLMs.
Today's best AI agents can only complete 33% of common online tasks like booking appointments or filling out job applications, revealing a significant gap between current capabilities and real-world utility.
A Qwen3-8B model, trained with a new SFT+RLAIF recipe on a challenging new benchmark, SWE-QA-Pro, beats GPT-4o in repository-level code understanding.