Search papers, labs, and topics across Lattice.
Tsinghua University's AI research group. Leading Chinese institution in NLP, knowledge graphs, and large language models.
100
0
0
PaDoc achieves a remarkable 67.4-118% increase in valid-page throughput while maintaining top-tier parsing accuracy, revolutionizing document parsing efficiency.
Smart-home agents struggle to differentiate between real commands and misleading ambient noise, with traditional detectors and MLLMs both failing in complementary ways.
AV-AIVAT enables agent evaluations to stop as soon as the evidence is sufficient, achieving a staggering 74x reduction in game requirements while maintaining statistical validity.
TrajDebug uncovers the root causes of failures in long-horizon agent trajectories, enabling targeted improvements that could significantly boost agent performance.
Operator-residual feedback slashes the rate of misleading score-only decisions from nearly 40% to under 2%, ensuring that autonomous agents make choices grounded in physical reality.
GeniWorld achieves robust zero-shot generalization in robotic manipulation, outperforming traditional models even with minimal training data.
Diff-Symbo achieves unprecedented quality and diversity in text-controlled music generation, outperforming leading models by leveraging a novel latent diffusion framework.
Event-adaptive compression allows EvtGraph to outperform traditional models while maintaining efficiency, proving that less can be more in temporal data representation.
Training LLM agents without expert supervision can lead to better performance and generalization across diverse environments.
Existing unlearning methods can leak sensitive knowledge through multi-hop reasoning paths, exposing a critical vulnerability in LLMs.
RTCF boosts the success of frozen VLA policies by leveraging past experiences without the need for retraining or extra GPU power.
Argus achieves a 78% success rate on long-horizon reasoning tasks while using 21% fewer tokens in mature workflows, showcasing a revolutionary approach to agentic autonomy.
AFD-Ledger reveals that optimizing deployment for AFD can drastically cut evaluation costs while exposing the nuanced performance dynamics between homogeneous and heterogeneous setups.
Constraint-First Reasoning reveals that explicitly managing answer-space constraints can dramatically enhance the accuracy of mathematical problem-solving in language models.
The proposed Attention-Guided Switching method allows MLLMs to dynamically balance between visual fidelity and logical coherence, achieving unprecedented efficiency in reasoning tasks.
SWD achieves high-fidelity circuit extraction with less than 1% of the data used by traditional methods, revolutionizing interpretability in pretrained transformers.
Language models exhibit a surprising bias towards cities with expansive infrastructure and rapid growth, revealing their implicit urban assumptions.
Dynamic adaptation in vision-language models can significantly boost performance while cutting down computational costs.
Image-level deepfake detectors can outperform traditional video-level detectors, with one achieving a remarkable 93.80% AUC when adapted for video analysis.
OmniPack achieves a remarkable 98% performance retention with a staggering 83.3% reduction in computational load, revolutionizing token compression for omni-modal models.
Accepting selective mismatches in autoregressive decoding can boost throughput by over 15% without any additional training or model adjustments.
DataSpace reveals that multimodal evidence integration can hinder accuracy in data agents, exposing critical gaps in current evaluation methods.
Synthesizing realistic 3D hand-object interactions from just a single image and text instruction could revolutionize AR/VR applications by enabling seamless integration of dynamic human behaviors.
CMIG-Net achieves up to a 0.619 dB gain in PSNR over existing methods by effectively leveraging conditional mutual information for low-light image enhancement.
Fourier decomposition of motion enables unprecedented accuracy in rendering dynamic scenes, tackling the limitations of traditional polynomial models.
SkillTrace redefines skill composition for LLM agents, achieving unprecedented success rates by leveraging a structured query-skill graph.
Tailoring deepfake detection to individual facial characteristics boosts accuracy and adaptability beyond traditional fixed architectures.
AdaThinkV achieves 40.79% accuracy in video reasoning while using 22.7% fewer tokens than its strongest adaptive baseline, showcasing a breakthrough in token-efficient reasoning.
TIDE-MC achieves unprecedented efficiency in matrix completion, handling billion-scale datasets without running out of memory while delivering up to 11,647x speedup.
Latent Softmax achieves up to 17.5% lower phoneme error rates in multilingual ASR by intelligently modeling tonal distinctions without sacrificing cross-lingual sharing.
Agents can now exploit flawed opponents safely, achieving up to 13.6 times the expected budget gain while certifying their own strategies.
Malicious audio instructions can stealthily hijack multimodal agents, achieving a 69.10% success rate in real-world scenarios.
Temporal reconstruction errors can be harnessed to significantly boost video quality, eliminating the need for expensive human annotations in preference optimization.
LedgerMind reveals that grounding multimodal reasoning in a structured evidence ledger can significantly mitigate common pitfalls like entity hallucination and unsupported reasoning.
ODEWorld achieves high-quality long-horizon predictions and planning-oriented dynamics by seamlessly integrating continuous-time modeling with ODE-based representations.
Agents can now autonomously teach themselves creative skills using high-quality human texts, bypassing the need for expensive human feedback.
CAM-DF reduces tool acquisition costs by 37% while maintaining task success, challenging the conventional wisdom that more tools always lead to better outcomes.
CURL effectively harnesses LLMs to stabilize CATE estimation, leading to improved performance in personalized interventions across multiple benchmarks.
UrbanDS outperforms traditional data science agents by effectively navigating complex urban datasets through a novel graph-guided multi-agent architecture.
Achieving up to 100% success on complex tasks without the need for search, INTACT redefines how we approach intent-to-action learning in dynamic environments.
Ranking systems can achieve greater accuracy and stability by leveraging structural relationships in permutations, as demonstrated by the SRPO framework.
Functionalization can lead to chemically distinct changes in electronic structure, revealing critical insights into lithium-metal electrolyte behavior.
Evolving solver and rubric skills in tandem reveals hidden weaknesses and boosts performance by up to 5% without relying on fixed evaluation criteria.
Achieving 86.9% accuracy in GUI task evaluation, the Interactive Reward Agent transforms how we assess and train GUI agents by integrating environment-state verification.
Steering signals are most effective when derived from execution-boundary states, not just from the presence of desired behaviors in the source text.
LaP-Forensics reveals that leveraging reconstruction-based evidence can significantly enhance deepfake detection accuracy against state-of-the-art generative models.
Despite high report quality, many models falter in citation accuracy and claim construction, exposing a disconnect between surface-level performance and deep reasoning skills.
CameraAnything enables filmmakers to reshoot videos with arbitrary camera angles and focal lengths in a single generation, revolutionizing video editing workflows.
High-curvature regions in point clouds can be effectively represented by a novel hyperbolic rectification method that boosts discriminative power and captures fine geometric details.
CADER redefines long-video reasoning by enabling systems to adaptively allocate resources based on confidence, significantly improving efficiency and accuracy.
A single policy label can mask significant differences in operational safety, with trusted-ledger strategies achieving over five times the authorized workflow completion compared to taint-only methods.
FCPAgent redefines how web agents validate their actions, achieving a 13.8% boost in success rates on complex tasks by integrating falsifiable commitments into planning.
Local interaction control in self-supervised denoising can dramatically enhance performance, revealing hidden regional dynamics that traditional methods overlook.
UNIFUSION achieves unprecedented performance in generative tasks by seamlessly adapting autoregressive models to uniform-noise diffusion, outperforming all evaluated models on key metrics.
BeyondFusion achieves high-quality infrared-visible fusion without the need for calibration, leveraging self-alignment to overcome sensor misalignment challenges.
RODR effectively disentangles optimization objectives to preserve geometric fidelity in point cloud denoising, overcoming a critical challenge in manifold representation.
MoNO achieves unprecedented per-prompt diversity in diffusion sampling while eliminating the need for auxiliary quality-control objectives.
Sketching at just 25% completion can yield better spatial rationality than fully specified baselines in 3D scene generation.
Achieving a 17.5x speedup in capacitance extraction while maintaining high accuracy could revolutionize electronic design automation workflows.
IRIS can detect model substitutions and routing dilutions in LLM gateways with unprecedented accuracy using only the output text, challenging the reliability of commercial AI services.
REFACT reduces token consumption while enhancing the density and faithfulness of reasoning traces in large language models, ensuring that every cited fact meaningfully supports the answer.
Lumera achieves state-of-the-art performance in 3D scene reconstruction, revealing critical gaps in light localization and cross-engine generalization.
Projecting raw scores onto the bridge polytope eliminates negative weights and boosts generative performance, leading to a remarkable reduction in perplexity.
Evaluator score gaps are the secret sauce for optimizing LLM policies, and DynamicRubric turns this insight into a powerful co-evolution framework that outshines existing methods.
Adaptive routing of perception priors allows PerceptDrive to generate optimal driving trajectories in real-time without complex post-processing.
OLEDLM can generate novel OLED candidates that meet stringent optoelectronic criteria, revolutionizing the search in a vast chemical space.
Evolving Cache Schedules can slash action-generation time by over 8x without sacrificing performance, revolutionizing real-time deployment of diffusion policies.
Vera achieves unprecedented identity consistency in human-centric video generation, drastically reducing identity confusion in multi-person scenarios.
Achieving high-fidelity face reconstruction in under 3 seconds, UVFaceFusion redefines the balance between speed and accuracy in digital avatar creation.
TAP-RAG achieves a remarkable +9.1 point accuracy boost over traditional multimodal RAG systems by tailoring evidence retrieval to specific query tasks.
MXSens redefines LLM quantization by achieving state-of-the-art accuracy with mixed precision, leveraging sensitivity-aware bitwidth allocation.
Adding depthwise convolutions to Transformers can boost accuracy on downstream tasks while barely increasing model size.
LLMs may ace rule-based tasks, but they falter in crucial reasoning areas like evidence integration and contradiction detection, revealing a significant gap in their utility for environmental law enforcement.
The Cramér-geometric Bellman operator reveals a unique fixed point that could transform how we approach evaluation errors in distributional reinforcement learning.
Revealing memory utilization patterns through attention can transform how agents refine their memory, leading to substantial gains in performance and efficiency.
Transitioning from 2D to 3D modeling reveals that fine details in monocular geometry can be captured with unprecedented fidelity.
Current LLMs struggle with multi-granularity event analysis, revealing critical performance gaps that could hinder their application in complex narrative tasks.
Achieving state-of-the-art multi-view hand-object interaction synthesis, HarmoHOI harmonizes 2D appearance and 3D motion in real-time.
OKR achieves a remarkable +6.5% mAP improvement over existing methods by preventing knowledge interference in domain-incremental object detection.
Rethinking GPU reliability, this study shows that ranking nodes by failure risk can outperform traditional predictive maintenance methods, capturing 64% of failures in the top 5% of at-risk GPUs.
Privacy risks in retrieval-augmented generation are not static; they vary dynamically with user queries, and our new framework addresses this critical oversight.
Pretraining MIL networks with knowledge distillation from foundation models boosts performance and stability, especially in few-shot scenarios.
Action QFormer boosts navigation success rates from 18.8% to 56.3% by intelligently reorganizing multimodal information under action supervision.
Achieving stable minute-long streaming and precise object manipulation in driving simulations could revolutionize how we approach autonomous vehicle training.
A unified framework reveals that the choice of tokenization and vocabulary topology can significantly influence the performance of discrete diffusion models, unlocking new avenues for optimization.
Continuous tracking of dynamic object evidence can transform MLLMs' ability to understand and interact with dynamic environments.
OrthoPilot outperformed seasoned orthopaedic experts in diagnostic reasoning, achieving a 10.6% increase in management success for complex musculoskeletal cases.
Personalized thumbnails can significantly boost user engagement by aligning visual content with individual preferences, outperforming traditional methods.
Motion4Motion revolutionizes motion transfer by eliminating the need for skeletons, allowing seamless animation across diverse species and character types.
TeleDexter achieves a remarkable 75% success rate in dexterous teleoperation tasks, where existing systems fail, showcasing a leap towards human-level control in robotic manipulation.
SAIL cuts cloud gaming bandwidth usage by over 44% while preserving perceptually lossless quality, revolutionizing cost-efficiency in the industry.
Automated synthesis can transform the personalization of animatronic faces, enabling rapid adaptation to diverse facial geometries with minimal manual intervention.
Training physics-informed neural networks with a unified priority framework can dramatically improve convergence and accuracy by respecting the physical information flow.
Temporal updates in LLMs can be made without sacrificing historical accuracy, achieving over 23% improvement in consistency with a single optimized representation.
A neuro-symbolic approach boosts the quality of twelve-tone music generation, elevating output consistency and expert preference significantly.
Adversarial attacks on vision-language agents reveal critical vulnerabilities, with multi-view optimization strategies proving significantly more effective than isolated approaches.
Achieving a 1.89× speedup in distributed neural rendering by cleverly minimizing synchronization barriers could redefine efficiency in large-scale training.
A structured evidence-state approach can boost omni-modal QA accuracy by over 30%, transforming how agents gather and validate information across diverse sources.
Generative retrieval can revolutionize statute retrieval by effectively bridging the gap between everyday legal language and formal statutes.
GTAlign achieves superior performance in graph classification tasks without the need for textual data, challenging the reliance on traditional graph neural networks and LLMs.