Search papers, labs, and topics across Lattice.
Fudan University
27
0
19
4
LLMs struggle with personal information retrieval, with the best model only achieving 57.3% accuracy on a new benchmark designed to evaluate mobile assistant capabilities.
Existing models mismanage tool use, but Beacon achieves a balance that enhances performance on complex tasks while preserving accuracy on simpler ones.
Change2Task recovers 29.2% more verified coding tasks than traditional methods, streamlining the training of coding agents.
Compounding error in video generation is linked to representational degradation, and a simple regularization technique can dramatically enhance long-term stability.
HiSkill bridges the gap between high-level skills and executable actions, enabling LLM agents to execute tasks more efficiently and effectively.
InterOCF reveals that effectively integrating 2D and 3D spatio-temporal dynamics can dramatically enhance the accuracy of occupancy forecasting for autonomous vehicles.
Existing HMR methods falter under severe occlusions, but the MVMP-HMR model sets a new standard for multiview multi-person mesh recovery in large-scale scenes.
Context-sensitive hallucinations in LLMs can be mitigated by a new optimization technique that rewards step-level consistency, leading to significant improvements in reasoning accuracy.
Separating action-driven dynamics from intrinsic world effects can boost planning success in complex environments by over 13%.
Subgoal-conditioned action generation boosts long-horizon planning success rates from 12.7% to 64.7% in complex environments.
A unified evaluation framework that simplifies the assessment of LLM-based agents could drastically enhance reproducibility and accelerate research breakthroughs.
Even state-of-the-art language models struggle significantly in real-world tasks, exposing critical shortcomings in their deployment readiness.
T2LDM++ generates realistic LiDAR scenes with rich geometric details, overcoming the limitations of existing models that struggle with insufficient training data and controllability.
JEPA-based world models face a fundamental trade-off between approximation and sample errors that could redefine their application in predictive tasks.
Even the top-performing language models struggle with archive-grounded reasoning, achieving only 59.4% accuracy on a benchmark designed to test their agentic capabilities across diverse workplace documents.
G2PO redefines agent actions and leverages a global state-transition graph, leading to a 22.2% boost in success rates for long-horizon tasks.
Transforming Poisson noise into Gaussian noise can boost image denoising performance by up to 0.75 dB, even in challenging conditions.
MemGUI-Agent achieves unprecedented long-horizon task performance by proactively managing context, outperforming traditional methods that struggle with prompt dilution.
Achieving lossless processing of 256K contexts, Keye-VL-2.0 transforms how we approach long-video understanding and agentic intelligence.
Superficial rephrasing can inflate AI peer review scores by over 1.3 points, revealing a dangerous vulnerability in AI-assisted scientific evaluation.
ISPO reduces critical reasoning failures in RLVR by transforming reward structures, leading to superior performance on complex reasoning tasks.
Surprisingly, the "think before answer" paradigm fails to enhance generative recommendation models, prompting a novel approach that redefines how reasoning is integrated into these systems.
SkillComposer enables language models to self-evolve skills in real-time, achieving up to +4.5 improvements on agent tasks compared to larger models.
LLMs' reliance on specific knowledge sources during question answering can now be reliably estimated without extensive perturbation, enabling better error detection and risk screening.
Naive distributed inference on edge devices can be *slower* than local execution due to CPU-GPU communication bottlenecks, but a profiling-driven adaptive approach can flip the script for significant gains.
Compressing 3D Gaussian splats by operating on intermediate feature representations slashes storage by an order of magnitude without sacrificing rendering quality.
RFT's impressive in-domain performance masks surprisingly weak generalization to new environments, highlighting a critical challenge for deploying LLM agents in the real world.