Search papers, labs, and topics across Lattice.

Alibaba's global research initiative. Publishes actively on NLP, multimodal models, and AI systems.
100
0
0
A single agentic framework can seamlessly tackle multiple intelligent document processing tasks, outperforming specialized models in the process.
EditBridge achieves ultra-high-resolution image editing at 4K with a speedup of up to 8.4 times, while maintaining high fidelity and perceptual quality.
Generative supervision can transform the landscape of moiré removal, yielding substantial performance gains even in the most challenging real-world scenarios.
LENS achieves 84.8% evidence recall without the overhead of pre-materializing evidence, redefining how we approach dynamic document retrieval.
CPI-Bench reveals significant performance gaps among image editing models, offering a more nuanced evaluation that aligns with real-world user experiences.
Realistic predictions in robotic manipulation now come with precise arm control, ensuring the right actions yield the expected outcomes.
HounsWorld achieves state-of-the-art performance in multimodal patient-state inference by seamlessly integrating CT imaging and clinical language, transforming how we interpret medical data.
Released tokenizer vocabularies can yield precise estimates of hidden corpus compositions, revealing insights that were previously obscured.
Emotion-driven feedback can transform multi-turn dialogue systems, boosting empathetic responses and model performance significantly.
Nearly 20% of LLM agent violations occur even after agents acknowledge safety constraints, highlighting a critical gap in execution awareness.
REST achieves few-step image generation that rivals traditional 40-step methods while slashing training costs by over 75%.
BCSD enables LLM agents to leverage external skills more effectively, achieving superior performance by integrating dual-context evaluations.
Task learnability can significantly enhance RL efficiency in LLMs, leading to better performance with less data.
C2C-Explorer boosts LLM inference efficiency, achieving a 44.1% increase in goodput and a staggering 98.4% reduction in memory usage.
BOUND achieves up to 5.6 EM points improvement over traditional methods by correcting local search-control errors in real-time.
FullDiT not only outperforms leading commercial music generators but also redefines how we approach music rendering by leveraging full-context generation techniques.
Event presence can mislead interpretations of skill execution fidelity, revealing critical discrepancies in agent performance assessments.
Real-time character animation is now feasible with Wan-Animate-2, which achieves high-fidelity results without the pitfalls of traditional motion representation methods.
Reducing token overhead by pruning redundant communication edges allows multi-agent systems to achieve better performance without the computational burden.
PrivDPO achieves robust LLM alignment while maintaining privacy, outperforming traditional methods in balancing privacy and utility.
Output extrapolation in STEP-OPD enables a unified student model to surpass the performance of its specialized teachers across all evaluated tasks.
SFC redefines semantic understanding in spoken language tasks, achieving superior accuracy and adaptability in open-domain contexts.
Current instruction-based video editing models are far from satisfactory, revealing critical gaps in evaluation that could reshape the field.
Token-level credit assignment in multi-turn language agents can be effectively enhanced by integrating teacher preferences through a novel self-distillation approach.
OmniPack achieves a remarkable 98% performance retention with a staggering 83.3% reduction in computational load, revolutionizing token compression for omni-modal models.
Sparse rewards can be effectively optimized without losing the advantages of fine-grained feedback through a novel two-stage training approach.
Solution Hacking reveals that up to 44.1% of answers from frontier LLMs may be misleadingly credited as correct due to invalid reasoning shortcuts.
Visual identity discrimination is the missing link in Universal Multimodal Embeddings, and this work provides a robust framework to integrate it seamlessly.
Follow-up suggestions for image editing can be made 32.70% more effective by integrating multimodal feedback into conversational systems.
AdaThinkV achieves 40.79% accuracy in video reasoning while using 22.7% fewer tokens than its strongest adaptive baseline, showcasing a breakthrough in token-efficient reasoning.
VC-Tooler achieves unprecedented adaptability in visual tool use, outperforming existing models by 60% in agentic reasoning tasks.
A new benchmark reveals that AI models can be rigorously evaluated on deep research tasks, highlighting their nuanced capabilities across diverse topics.
Misalignment between visual evidence and predicted timestamps in video grounding can lead to substantial performance drops, but CAVE effectively bridges this gap with boundary-specific rewards.
LoopMemGR transforms generative recommendation by integrating past recommendations into the memory framework, leading to more informed and effective predictions.
Synthesized egocentric videos can elevate real-robot task success rates by over 10% through enhanced training data.
ReAlloc not only optimizes marketing budget allocation but also achieves simultaneous lifts in pay order and income, challenging traditional approaches in e-commerce.
LLMs can identify flagged security issues but struggle to uncover silent intrusions and create effective remediation plans, revealing a critical gap in their utility for real-world incident response.
SERPO achieves up to 20.63 points improvement on HealthBench by enabling language models to self-evolve their evaluation criteria in real-time without external feedback.
Evolving solver and rubric skills in tandem reveals hidden weaknesses and boosts performance by up to 5% without relying on fixed evaluation criteria.
OmniDelta achieves a 1.64x speedup in inference while reducing GPU memory usage by 22% at just 25% token retention, redefining efficiency in OmniLLMs.
TopoGR reveals that preserving the topology of semantic IDs can dramatically enhance item relatedness in generative recommendation systems, leading to superior performance.
MTGuard effectively reduces harmful tool use in LLM agents, combining static and dynamic analysis to enhance security without sacrificing performance.
SWAG-Bid redefines auto-bidding by integrating long-term market forecasting and adaptive control, leading to superior advertising outcomes.
Context-aware routing and compression in SPARC leads to superior generative recommendations by preserving crucial information without inflating input size.
Achieving a 60x speedup in avatar generation without compromising visual quality could revolutionize production workflows in multimedia content creation.
Achieving superior auto-bidding performance with less than 10% of the fine-tuning parameters required by traditional methods could revolutionize how LLMs are applied in dynamic auction systems.
TreeAdapter revolutionizes species image generation by harnessing hierarchical taxonomic structures, achieving unprecedented detail and accuracy in fine-grained outputs.
Repeated revisions in coding agents can lead to a significant drop in reliability, highlighting the need for structured evidence-bound contracts in code repair processes.
Achieving real-time audio-video generation without sacrificing visual fidelity or synchronization, TaoMate redefines the limits of digital human technology.
OmniScope reveals that treating audio and video relevance separately can drastically enhance performance in omnimodal models, achieving remarkable efficiency gains without sacrificing accuracy.
VLM-IE3D achieves state-of-the-art performance in 3D tasks by seamlessly integrating implicit and explicit geometric representations from RGB inputs.
Current SID-based generative recommendation models can sometimes reach future items but falter with completely unseen tokens, revealing critical gaps in cold-start handling.
ECoM Reasoning boosts spoken language model accuracy by 21% while slashing token usage to just 40% of the standard approach.
Achieving high visual fidelity while ensuring physical consistency in video generation could redefine the standards for simulating realistic interactions in AI-generated content.
Open-Vocabulary Gaze Object Prediction can now effectively recognize and localize objects from a diverse array of unseen categories, outperforming traditional methods.
TSGR redefines e-commerce search by ensuring that high-value items are prioritized in retrieval, achieving measurable business impact in a competitive market.
Achieving up to 17× faster inference for long-context LLMs without compromising output quality could redefine efficiency standards in AI applications.
CODA redefines edge video diffusion by achieving up to 1.80x faster inference and 1.74x better energy efficiency without sacrificing output quality.
TmallGS redefines e-commerce search ranking by optimizing feature representation and interaction, resulting in substantial performance gains over traditional models.
Aesthetic assessments can be dramatically improved by focusing on key moments and endings, as shown by Peak-End-Net's state-of-the-art results in video evaluation.
Over 60% reduction in post-purchase redundancy reveals a critical flaw in how traditional recommendation systems interpret user intent.
Incorporating discount rates into conversion predictions can boost online sales performance by over 3% in real-world applications.
Bias in LLMs can be traced to specific geometric structures in their hidden states, allowing for precise control over scoring outcomes.
A structured evidence-state approach can boost omni-modal QA accuracy by over 30%, transforming how agents gather and validate information across diverse sources.
AWA-RL boosts precision in search agents by up to 10.3% by penalizing hallucinations, challenging the status quo of LLM training.
Thinking Collapse can severely impair reasoning in LLMs, but a new adaptive framework boosts accuracy by over 4% while preserving cognitive capacity.
Balancing session-centric scheduling can boost LLM cluster throughput by up to 16% without sacrificing token reuse.
GeoProp achieves a remarkable 10.6% boost in real-world manipulation tasks by effectively grounding robot state in visual context, all while adding minimal complexity.
SAYRE's innovative approach to synthesizing KIE training data leads to substantial performance gains for on-device models, particularly in challenging extraction scenarios.
CanniUplift boosts e-commerce platform growth by over 4% by tackling the dual challenges of seller and incentive cannibalization in uplift modeling.
Cold start latencies can be slashed by up to 99.3% with CoCoScale's innovative layer-wise scaling approach.
UniSGR achieves a breakthrough in recommendation systems by seamlessly integrating semantic ID generation with multi-objective ranking, leading to superior performance in e-commerce applications.
OPSD can backfire, leading to rote memorization instead of enhancing reasoning, but a novel decomposition approach reveals a path to meaningful improvements.
40% to 73% of multi-turn coding tasks lose previously correct behavior, highlighting a critical flaw in LLM-assisted software development.
HEE transforms static image understanding into a dynamic, query-guided exploration that significantly boosts accuracy in high-resolution perception tasks.
Memory overload in autoregressive video generation can be tackled by absorbing historical context into model weights, achieving up to 50% cache reduction with minimal quality loss.
GUI agents can now leverage an actively maintained memory state to significantly improve task execution accuracy and efficiency over long horizons.
AtomiMed achieves a strikingly higher correlation with human radiologist evaluations, transforming how we assess clinical report accuracy.
LLMs struggle with compositional reasoning, showing sharp performance drops in executing order-sensitive data refinement tasks.
Current interactive world models fall short, with none passing the rigorous tests of WorldRoamBench designed to assess long-horizon stability across action, vision, physics, and memory.
The Relative Surprisal Index reveals that the interplay between token probability and entropy is crucial for optimizing reinforcement learning in language models, leading to substantial performance gains.
CIPE-Dance is a game-changer, providing the largest dataset for dance video generation and enabling OmniDance to set new benchmarks in multimodal video synthesis.
Dynamo achieves a remarkable 5.6% average accuracy boost in visual reasoning tasks without retraining, revolutionizing how VLMs adapt in real-time.
UniGP reveals that joint training of controllable generation and dense prediction can significantly enhance performance without the need for complex designs, outperforming specialized models.
InnerZoom cuts end-to-end latency by up to 31.8% while surpassing existing GUI grounding methods, proving that less can indeed be more.
Evaluating LLM agents in microservice failure diagnosis reveals that traditional outcome-based benchmarks miss critical reasoning processes, which these new datasets effectively capture.
By harnessing clean latent features, SharpMoE dramatically improves resource allocation to salient tokens, overcoming a key limitation in diffusion MoE architectures.
CHAUN achieves a remarkable 25.6% improvement in QINI scores, redefining the benchmarks for uplift modeling in the presence of unobserved confounding.
AIGP not only boosts e-commerce profitability by over 13% but also delivers interpretable pricing strategies that align with long-term business goals.
Language-action pretraining can lead to VLA policies that are not only more robust but also less dependent on visual cues, achieving up to 45% higher success rates in real-world tasks.
Selected features from sparse autoencoders can causally steer language models toward desired behaviors, like refusal, revealing new avenues for interpretability and control.
A tailored vision-language framework boosts diagnostic accuracy on 3D CT imaging by over 5% while enhancing zero-shot reliability through innovative semantic alignment techniques.
PolicyAlign enables LLMs to adapt to rapidly changing safety policies without relying on expensive supervision data, achieving significant safety improvements across diverse applications.
SARA unlocks the potential of low-resource languages in multilingual models by aligning their expert routing with high-resource anchors, leading to measurable performance gains.
REVERIEMEM boosts character fidelity in role-playing agents, achieving a staggering 34.6% improvement in knowledge accuracy while preserving unique character voices.
TryOnCrafter revolutionizes video virtual try-on by enabling dynamic, omnidirectional viewpoint exploration without sacrificing realism or structural integrity.
EchoStyle achieves high-quality video stylization without the content leakage and style drift that plague existing methods, even outperforming many proprietary systems.
HDS achieves 44% faster training iterations while improving model performance, redefining efficiency in LLM pre-training.
STAR-VAE achieves state-of-the-art audio reconstruction fidelity by aligning latent space geometry with the hierarchical structure of audio signals.
AudioCALM achieves state-of-the-art performance in speech, sound, and music generation by seamlessly integrating diverse audio modalities into a single autoregressive framework.