Search papers, labs, and topics across Lattice.
Tsinghua University's AI research group. Leading Chinese institution in NLP, knowledge graphs, and large language models.
100
7
0
UPT can either refine model capabilities or amplify errors, depending on the internal signals used during adaptation.
Retaining future imagination through compact latent actions allows LAWA to outperform existing models while slashing inference latency by nearly 43%.
OPDVR transforms the landscape of model distillation by ensuring that only correct trajectories enhance learning, leading to significant performance gains on reasoning tasks.
DIAG reshapes practice distribution to maximize informative supervision, leading to significantly improved reasoning performance in LLMs.
Object-Uni transforms how we understand and generate spatial representations of objects, enabling precise manipulation of their poses in generated images.
TRACE transforms high-performing LLMs into consistently reliable agents, achieving a remarkable 34.6-point boost in task consistency.
Current scientific agents struggle to maintain a coherent narrative across evidence and calculations, with only 34.81% achieving strict accuracy in complex tasks.
LLM agents struggle with network configuration, revealing failures that extend beyond simple command errors to deeper issues in task adherence and planning.
A decentralized bidding system for LLM agents not only enhances efficiency but also reduces manipulation risks, outperforming traditional orchestration methods.
Agents using ParallelWorld can efficiently evaluate multiple future trajectories, leading to superior decision-making in complex environments.
Achieving a 25% reduction in prediction error, MOSH-WM redefines the landscape of object-centric video forecasting by grounding state representations in visual support.
Jiuge-Tuiqiao transforms AI from a passive generator into an active collaborator, enhancing user creativity in classical Chinese poetry.
A deep learning model predicts a significant summer drought in central China for 2026, driven by atmospheric circulation patterns linked to Pacific warming.
By making environment design a learnable process, SPADE unlocks a new frontier in self-improvement for language agents, leading to substantial performance gains across diverse tasks.
Scalar metrics fail to capture the true diversity of AI-generated content, but diversity profiles offer a robust, multi-dimensional evaluation framework that reveals hidden biases.
GeoWeaver achieves unprecedented accuracy in long-sequence 3D reconstruction by correcting accumulated errors through a novel combination of geometric priors and adaptive refinement techniques.
MSEditor achieves unprecedented consistency in multi-shot video editing, outperforming existing methods by effectively managing identity drift and temporal coherence.
Traffic element awareness can dramatically elevate the performance of autonomous driving systems, achieving state-of-the-art results with minimal architectural changes.
Separating spatial and temporal modeling in action recognition leads to significant performance gains, challenging the effectiveness of existing implicit coupling methods.
Achieving the best macro-average results on dynamic scene reconstruction benchmarks, UniQuery4R redefines efficiency in 4D scene understanding.
Simon's algorithm may not be as robust as previously thought, with critical lemmas lacking sufficient proof to ensure its correctness.
Domain adaptation can dramatically enhance sign language recognition, achieving superior results over conventional transfer learning methods.
Guiding dexterous grasp generation with arm-aware constraints boosts feasibility in complex environments, outperforming traditional methods reliant on hand-centric models.
The form of answer labels, not just their quantity, fundamentally shapes what LLMs learn during fine-tuning, revealing a surprising causal relationship that could redefine training strategies.
Achieving 66.9 MOTA in 3D tracking without LiDAR reveals a new frontier in roadside infrastructure understanding.
Closed-loop learning in embodied agents can lead to a staggering 11.1x speedup in inference while achieving record performance on complex tasks.
The empirical analysis reveals that larger models may excel in generating narratives but often fail to maintain coherence and depth, exposing a critical trade-off in LLM storytelling capabilities.
Vision-based tactile sensors could revolutionize robotic interaction by providing high-resolution tactile data that enhances perception and manipulation capabilities.
CineDub achieves unprecedented accuracy in multi-speaker dialogue dubbing from uncropped videos, setting a new standard for video dubbing technology.
Surprisingly, larger LLMs benefit from increased repetition of high-quality domain data, challenging conventional wisdom about data diversity in training.
ROLoad-PMP achieves robust security for low-level software with less than 1.40% hardware overhead while enabling lightweight defenses that outperform existing solutions.
Visual perturbations can significantly alter predictions in world models, but ACPC offers a quantifiable way to diagnose and mitigate these effects.
Realistic predictions in robotic manipulation now come with precise arm control, ensuring the right actions yield the expected outcomes.
SCOUT achieves a remarkable 16.85% improvement in spatial reasoning benchmarks, setting a new standard for Vision-Language Models.
TideRL boosts RL training goodput by up to 5.6 times, transforming how we approach efficiency in multi-turn agentic workloads.
Cross-path reasoning reveals untapped research ideas, outperforming traditional methods and rivaling human insights in scholarly evolution.
Released tokenizer vocabularies can yield precise estimates of hidden corpus compositions, revealing insights that were previously obscured.
Auditing Chinese web content reveals pervasive pollution that shifts over time, challenging the integrity of LLM training data.
ChemWorld allows researchers to isolate the impact of hidden chemical laws on agent behavior, enabling unprecedented control and replayability in autonomous chemistry experiments.
Verifiable temporal grounding in video forensics can drastically improve the detection of AI-generated content, outperforming traditional methods reliant on coarse supervision.
Pruning 50% of channels in RGB-infrared object detectors can actually boost performance by 0.6% mAP, challenging conventional wisdom about redundancy.
Cross-frame feedback boosts Transformer tracking performance by leveraging historical information, outperforming traditional same-frame methods by up to 3.2 AO points.
Even the top-performing LLM struggles with complex legal temporal reasoning, revealing significant gaps in AI's understanding of time-sensitive legal contexts.
Simulation traces can transform LLMs into powerful tools for diagnosing and improving complex scheduling policies, achieving unprecedented performance gains.
Revisiting the same locations across time in Sekai2 enables the learning of persistent scene representations, a game-changer for interactive world modeling.
Understanding how regularization influences model learning could revolutionize our approach to designing robust AI systems.
GameAlpha-2.4K not only revolutionizes RGBA video generation for gaming but also achieves a significant efficiency boost by intelligently bypassing redundant computations.
A physics-guided approach in VeinCast reveals that integrating structured atmospheric relations can significantly enhance the accuracy of medium-range weather forecasts.
Integration of diverse robot policies can be streamlined from hours to minutes, revolutionizing how we deploy and evaluate robotic systems.
Achieving state-of-the-art performance in both 4D reconstruction and point tracking, Uni4R leverages continuous velocity fields to model dynamics at any timestamp, breaking free from traditional limitations.
Unpaired MRI translation can effectively preserve anatomical structures across varying magnetic field strengths, a significant leap in addressing domain shift challenges in medical imaging.
BAMU achieves a remarkable MOS improvement in speech quality by dynamically allocating quantization resources based on frame complexity, outperforming traditional fixed-depth codecs.
OnEvoMemory allows robots to evolve their memory in real-time, leading to improved task performance and reduced redundancy in long-horizon manipulation tasks.
Even the most advanced LLMs struggle to maintain narrative consistency, with a staggering 68% of generated content conflicting under user interventions.
CED reveals that VLMs can be trained to prioritize evidence-based reasoning over language shortcuts, leading to more reliable visual understanding.
A multi-agent forensic reasoning framework outperforms leading closed-source models in deepfake detection by leveraging diverse analytical perspectives on forgery cues.
Operator-residual feedback slashes the rate of misleading score-only decisions from nearly 40% to under 2%, ensuring that autonomous agents make choices grounded in physical reality.
TrajDebug uncovers the root causes of failures in long-horizon agent trajectories, enabling targeted improvements that could significantly boost agent performance.
AV-AIVAT enables agent evaluations to stop as soon as the evidence is sufficient, achieving a staggering 74x reduction in game requirements while maintaining statistical validity.
GeniWorld achieves robust zero-shot generalization in robotic manipulation, outperforming traditional models even with minimal training data.
Smart-home agents struggle to differentiate between real commands and misleading ambient noise, with traditional detectors and MLLMs both failing in complementary ways.
Trajectory scoring in aerial navigation can be revolutionized by focusing on unexplainable prediction discrepancies, leading to more robust and efficient UAV navigation.
Diff-Symbo achieves unprecedented quality and diversity in text-controlled music generation, outperforming leading models by leveraging a novel latent diffusion framework.
Existing unlearning methods can leak sensitive knowledge through multi-hop reasoning paths, exposing a critical vulnerability in LLMs.
Training LLM agents without expert supervision can lead to better performance and generalization across diverse environments.
Hard prompt compressors can leave critical context gaps, leading to a staggering 60% of examples suffering from referential dangling, which severely impacts accuracy in multi-hop question answering.
Constraint-First Reasoning reveals that explicitly managing answer-space constraints can dramatically enhance the accuracy of mathematical problem-solving in language models.
AFD-Ledger reveals that optimizing deployment for AFD can drastically cut evaluation costs while exposing the nuanced performance dynamics between homogeneous and heterogeneous setups.
RTCF boosts the success of frozen VLA policies by leveraging past experiences without the need for retraining or extra GPU power.
Argus achieves a 78% success rate on long-horizon reasoning tasks while using 21% fewer tokens in mature workflows, showcasing a revolutionary approach to agentic autonomy.
Event-adaptive compression allows EvtGraph to outperform traditional models while maintaining efficiency, proving that less can be more in temporal data representation.
Language models exhibit a surprising bias towards cities with expansive infrastructure and rapid growth, revealing their implicit urban assumptions.
Dynamic adaptation in vision-language models can significantly boost performance while cutting down computational costs.
Accepting selective mismatches in autoregressive decoding can boost throughput by over 15% without any additional training or model adjustments.
Image-level deepfake detectors can outperform traditional video-level detectors, with one achieving a remarkable 93.80% AUC when adapted for video analysis.
SWD achieves high-fidelity circuit extraction with less than 1% of the data used by traditional methods, revolutionizing interpretability in pretrained transformers.
OmniPack achieves a remarkable 98% performance retention with a staggering 83.3% reduction in computational load, revolutionizing token compression for omni-modal models.
The proposed Attention-Guided Switching method allows MLLMs to dynamically balance between visual fidelity and logical coherence, achieving unprecedented efficiency in reasoning tasks.
DataSpace reveals that even the best multimodal models struggle with data agent accuracy, achieving only 66.34% in complex heterogeneous environments.
CMIG-Net achieves up to a 0.619 dB gain in PSNR over existing methods by effectively leveraging conditional mutual information for low-light image enhancement.
Tailoring deepfake detection to individual facial characteristics boosts accuracy and adaptability beyond traditional fixed architectures.
Synthesizing realistic 3D hand-object interactions from just a single image and text instruction could revolutionize AR/VR applications by enabling seamless integration of dynamic human behaviors.
Fourier decomposition of motion enables unprecedented accuracy in rendering dynamic scenes, tackling the limitations of traditional polynomial models.
AdaThinkV achieves 40.79% accuracy in video reasoning while using 22.7% fewer tokens than its strongest adaptive baseline, showcasing a breakthrough in token-efficient reasoning.
SkillTrace redefines skill composition for LLM agents, achieving unprecedented success rates by leveraging a structured query-skill graph.
TIDE-MC achieves unprecedented efficiency in matrix completion, handling billion-scale datasets without running out of memory while delivering up to 11,647x speedup.
Latent Softmax achieves up to 17.5% lower phoneme error rates in multilingual ASR by intelligently modeling tonal distinctions without sacrificing cross-lingual sharing.
Agents can now exploit flawed opponents safely, achieving up to 13.6 times the expected budget gain while certifying their own strategies.
Malicious audio instructions can stealthily hijack multimodal agents, achieving a 69.10% success rate in real-world scenarios.
LedgerMind reveals that grounding multimodal reasoning in a structured evidence ledger can significantly mitigate common pitfalls like entity hallucination and unsupported reasoning.
Agents can now autonomously teach themselves creative skills using high-quality human texts, bypassing the need for expensive human feedback.
ODEWorld achieves high-quality long-horizon predictions and planning-oriented dynamics by seamlessly integrating continuous-time modeling with ODE-based representations.
Temporal reconstruction errors can be harnessed to significantly boost video quality, eliminating the need for expensive human annotations in preference optimization.
UrbanDS outperforms traditional data science agents by effectively navigating complex urban datasets through a novel graph-guided multi-agent architecture.
CURL effectively harnesses LLMs to stabilize CATE estimation, leading to improved performance in personalized interventions across multiple benchmarks.
CAM-DF reduces tool acquisition costs by 37% while maintaining task success, challenging the conventional wisdom that more tools always lead to better outcomes.
Ranking systems can achieve greater accuracy and stability by leveraging structural relationships in permutations, as demonstrated by the SRPO framework.
Evolving solver and rubric skills in tandem reveals hidden weaknesses and boosts performance by up to 5% without relying on fixed evaluation criteria.
Steering signals are most effective when derived from execution-boundary states, not just from the presence of desired behaviors in the source text.
Achieving 86.9% accuracy in GUI task evaluation, the Interactive Reward Agent transforms how we assess and train GUI agents by integrating environment-state verification.