Search papers, labs, and topics across Lattice.
One of the world's largest corporate research labs, spanning AI, systems, and human-computer interaction.
100
0
0
Fast-weight updates in LongVU-TTT enable MLLMs to retain crucial visual context, significantly boosting performance in long video understanding tasks.
Trace Integrity reveals that LLMs can produce seemingly correct answers backed by invalid computations, challenging the reliability of traditional evaluation metrics.
Active learning can dramatically reduce the annotation burden in summarization tasks, with LOBSTER achieving up to 665x faster query selection without sacrificing performance.
Slasher can dynamically adjust Azure datacenter power consumption with minimal impact on workloads, addressing critical energy management challenges in cloud computing.
pigzpp achieves up to 16 times faster compression than Python's gzip while maintaining full compatibility with existing gzip standards.
BALIGN filters out high-risk preference samples, preserving foundational model capabilities while optimizing alignment, achieving the best of both worlds.
Current vision-language models falter in 3D assembly reasoning, struggling with accuracy and complexity in real-world industrial scenarios.
A deep learning model predicts a significant summer drought in central China for 2026, driven by atmospheric circulation patterns linked to Pacific warming.
Segments with conflicting quality signals are systematically undervalued, revealing critical flaws in current evaluation metrics for multilingual tasks.
KNOWSIM uncovers that LLM performance is not uniform and shifts dramatically based on user knowledge levels, challenging traditional evaluation methods.
Dion3 slashes optimizer step time by up to 6x while maintaining or improving loss performance compared to its predecessor, Muon.
Colorization methods can now adaptively handle both modern and historical grayscale images, reducing color artifacts and improving visual fidelity.
Abstaining from uncertain predictions can enhance LLM accuracy in medical applications by 9.6 percentage points, transforming uncertainty into a strategic advantage.
OEO shows that a capable optimizer can outperform traditional pipelines, achieving 12 wins in 14 comparisons while using significantly fewer resources.
Eco-SoC achieves a remarkable 42% reduction in switching activity while offsetting its carbon footprint in just over a year of deployment.
Self-distillation conditioned on privileged information may lead to a model that is less capable of reasoning, as it optimizes for a misleading signal rather than task success.
Self-evolving rubric rewards can dramatically enhance audio reasoning in models, outperforming traditional methods by adapting to the model's evolving capabilities.
Evaluation criteria in LLM-as-judge systems are often interdependent, and RADAR reveals these hidden couplings that can skew decision-making processes.
Causal reasoning reveals that specific prompt designs can detrimentally affect code generation accuracy in LLMs, challenging assumptions about optimal input strategies.
A novel curation pipeline can optimize regression evaluation sets by maximizing capability coverage while adhering to strict query limits.
Evaluating AI agents in a dynamic enterprise environment reveals that static snapshots miss critical context, but a new system allows for real-time, persona-driven assessments of agent performance across time.
LLM agents fabricate product attributes in over half of their listings, but a novel reputation-penalty mechanism can significantly curb this behavior without needing access to the truth.
The rise of LLM agents threatens to erase the clarity of authorship and accountability in collaborative knowledge work, raising urgent questions about intellectual integrity.
Evaluators using cESA can achieve more reliable translation quality assessments while cutting annotation time by leveraging shared context across multiple outputs.
Distilling from weaker models can enable a student to outperform its stronger counterparts, challenging the conventional wisdom of model hierarchy in AI training.
DVPSFormer reduces the computational burden of depth-aware video panoptic segmentation, enabling real-time decision-making for autonomous vehicles without sacrificing accuracy.
Specula uncovers deep bugs in system code that traditional methods often miss, revolutionizing formal specification generation.
Response-only moderation maximizes usefulness, but combining input and response strategies can significantly reduce harmful exposure—revealing a critical trade-off in content moderation design.
Modular robot policies synthesized through natural language corrections outperform traditional black-box models, enabling interpretable and adaptable robotic behaviors.
OpenForgeRL allows researchers to train AI agents in real-world environments with unprecedented ease and efficiency, revealing that some harnesses are significantly harder to learn than others.
DMG achieves a remarkable 4.9X performance boost while slashing compute-side cache requirements by nearly 19X, redefining efficiency in graph processing systems.
State-of-the-art LLMs fail to capture nuanced user preferences, lagging behind simple baselines in predicting choices in interactive narratives.
PRTA outperforms traditional and LLM-based recommendation systems by effectively leveraging multiple models through a central LLM planner, enhancing personalization without the pitfalls of hallucination.
Lax bug reproduction tests can lead to plausible but incorrect patches, but a new iterative framework boosts repair success by refining both tests and fixes.
Agentic reasoning tools struggle with complex financial documents, revealing substantial gaps in their capabilities that could impact decision-making in finance.
ATLAS achieves over 500-fold efficiency in sampling amorphous materials while maintaining less than 0.2% free energy error, revolutionizing the approach to material design.
Operators can now dynamically analyze distributed traces across multiple dimensions, enhancing anomaly diagnosis with tailored insights from both visual and natural language interfaces.
Real-world deployment of LLM-based agents reveals critical safety and reliability challenges that traditional benchmarks overlook.
Transitioning from 2D to 3D modeling reveals that fine details in monocular geometry can be captured with unprecedented fidelity.
ReViV reconstructs 4D viewer and view dynamics from a single monocular video, achieving unprecedented accuracy and speed without heavy task-specific priors.
SciForma achieves unprecedented structural fidelity in scientific diagrams, outperforming both open-source and proprietary models.
Natural paper revisions can be harnessed to train AI agents for precise and context-aware editing of complex scientific diagrams.
TRACE transforms how long-horizon agents are trained, leading to a remarkable performance increase on complex tasks without the need for supervised fine-tuning.
Encoded adversarial prompts can exploit AI safety mechanisms with alarming effectiveness, revealing significant vulnerabilities in current models.
HASTE enables rapid, accurate building damage assessment in disaster zones, delivering actionable insights within hours using minimal user input.
Data synergy can either amplify or diminish model performance, revealing that the right dataset combinations are crucial for optimal language model training.
Current MLLMs miss critical insights from multi-view sports videos, but a new agentic framework boosts performance by over 14% through smarter view selection.
Transforming isolated memory pools into a cohesive resource can reduce cache miss rates by up to 63% for memory-constrained tenants.
Current LLMs only achieve 27.3% accuracy in reasoning about scientific lineage, revealing a critical gap in their compositional capabilities.
Layer patching can dramatically enhance model performance in size interpolation, revealing that simple strategies often outperform complex methods.
MedPMC transforms millions of clinical articles into a goldmine of high-fidelity image-text pairs, dramatically improving model performance and clinical relevance.
TurnOPD redefines on-policy distillation by optimizing training budgets at the turn level, leading to superior agent performance without increasing training time.
CAIRN redefines 3D scene understanding by seamlessly integrating room-level topology with object-level relations, achieving unprecedented performance in multi-room environments.
LangLoc achieves unprecedented accuracy in indoor localization from natural language, closing the gap between coarse scene retrieval and precise pose estimation.
Targeted feedback can slash calculation errors in small language models from 56.9% to 23.5%, revolutionizing their physics reasoning abilities.
$λ$-VAE achieves up to 2.8x more information capacity while preventing posterior collapse in VAEs through a novel variance equalization technique.
Temporal domain adaptation can dramatically enhance high-resolution climate projections, especially in challenging topographical regions.
ResearchStudio-Idea transforms the ideation process by systematically grounding proposals in literature and identifying unresolved research bottlenecks, leading to more robust and traceable research directions.
ResearchStudio-Reel not only automates research dissemination but does so with unprecedented quality, outperforming both traditional methods and leading LLMs in aesthetic appeal and information accuracy.
Code LLMs can recognize incorrect instructions but still follow them, leading to irrecoverable semantic errors that defy traditional evaluation metrics.
Interleaving speech and text during ASR training boosts entity recognition accuracy and narrows the gap between modalities, challenging traditional training paradigms.
Prefill-deflecting scheduling can cut Time-to-First-Token by up to 81%, revolutionizing disaggregated LLM serving efficiency.
Ink3D achieves a breakthrough in 3D asset creation, enabling the generation of complex textures that were previously unattainable with conventional methods.
Developers are more likely to trust AI with decision-making in high-demand tasks, but resist autonomy in work that defines their professional identity.
ELDR slashes median latency by up to 13.9% for MoE models by intelligently routing requests based on expert activation signatures.
By treating slide design as an inverse planning problem, SPIRE reveals latent design intents that traditional methods miss, leading to superior personalization outcomes.
Achieving up to 3.5% mAP gains and 1.3x higher throughput, RT-SFOD redefines the trade-offs in source-free object detection by being faster and more compact without sacrificing accuracy.
Regret bounds that defy traditional scaling laws could revolutionize how we approach contextual slate bandit problems in adaptive settings.
LOTUS achieves a groundbreaking 2.5x-6.9x reduction in reasoning latency while matching explicit chain-of-thought performance at 3B parameters.
PRIME-Speech achieves low-latency, accurate speech-to-speech generation without sacrificing the robust performance of existing speech-to-text models.
Mandol achieves a 5.4x speedup in retrieval and a 4.8x speedup in insertion, revolutionizing long-term conversational memory management.
Failed rollouts can be a goldmine for training, revealing insights that lead to significant performance improvements in zero-hit reasoning scenarios.
TF-MoE achieves a remarkable +3.8 dB improvement in speech separation performance while keeping computational costs low, making it ideal for edge-device applications.
Reusable procedural skills derived from agent traces can drastically cut down execution time and boost success rates in complex tasks.
Readers find AI-generated translations "fine," but overwhelmingly prefer human translations for their clarity and immersive quality, despite being unable to reliably distinguish between the two.
DeformGen transforms the landscape of deformable manipulation by enabling effective policy learning through innovative state augmentation and trajectory adaptation techniques.
Clarifying memories can significantly boost the factual accuracy and personalization of conversational agents, while irrelevant memories lead to degraded responses.
Low-bit quantization can inflate reasoning length, leading to hidden compute costs that traditional accuracy metrics overlook.
Achieving comparable performance to full-precision models, BITEMBED slashes storage costs and enhances embedding efficiency with extreme low-bit quantization.
ConflictScore reveals that language models often overlook conflicting evidence, leading to overconfident and inaccurate claims.
Spurious correlations in foundation models can be effectively disentangled using a dual-branch approach, achieving superior bias mitigation with minimal parameter adjustments.
MambaRaw achieves a remarkable 1.4 dB increase in PSNR at low metadata bitrates while slashing coding latency by nearly 9%, setting a new benchmark in raw image reconstruction.
Mistakes in human demonstrations can enhance robot learning when properly harnessed, revealing a new dimension of value estimation that traditional methods overlook.
Asynchronous OPD can boost training throughput significantly while managing the challenges of stale data, transforming the efficiency of large language model fine-tuning.
IRENE not only enhances zero-shot retrieval accuracy but also drives a 4.2% increase in ad click-through rates in live environments.
D2D transforms conversational product search by cutting conversation times by nearly 30% while boosting accuracy and user satisfaction.
High-diversity training improves safety in VLA models, but sub-optimal trajectory synthesis still hinders task success.
Just two factors can explain over 90% of a model's performance across 133 benchmarks, drastically simplifying evaluation processes.
Evaluation awareness in language models reveals a significant gap between benchmark performance and real-world safety, challenging the reliability of current evaluation metrics.
Agentic AI could revolutionize cybersecurity by transforming labor-intensive tasks into efficient, automated defenses.
G2PO redefines agent actions and leverages a global state-transition graph, leading to a 22.2% boost in success rates for long-horizon tasks.
Prospective memory in LLMs is not just harder than retrospective memory; it reveals critical insights into a model's reasoning capacity and attentional robustness.
Off-policy degree in RLVR updates can drastically change which tokens drive learning, leading to a new adaptive method that outperforms traditional baselines.
Disciplinary siloing in research is starkly revealed through a novel citation graph that links claims to their sources, reshaping our understanding of knowledge evolution in AI fields.
Cognitive diversity among developers leads to distinct interaction modes with programming assistants, revealing that one-size-fits-all solutions may fall short.
Achieving 60 FPS dynamic 4D hand reconstruction from egocentric videos, Hand-4DGS outperforms traditional methods by effectively handling occlusions and rapid motion.
MuseVLA achieves an impressive 80.6% success rate in robotic manipulation tasks by leveraging diverse sensing modalities, surpassing traditional RGB-only models.
Tail latency in LLM serving can be cut by up to 50% without relying on length predictions, reshaping how we optimize inference performance.
CoTriSyGen achieves unprecedented long-range coherence in video generation by integrating visual evidence into a dynamic memory system, drastically reducing identity drift across shots.
FastContext cuts coding agent token usage by 60% while boosting resolution rates by 5.5% by decoupling code exploration from task-solving.