Search papers, labs, and topics across Lattice.

Top-tier US AI research university. Strong in NLP, ML systems, and computer vision.
100
0
0
Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.
Multimodal agents do not need external verifier modules to handle noisy retrieval—targeted RL rewards can teach a 7B model to internally audit search results and approach proprietary-tier multi-hop reasoning on just 5,000 training samples.
Moving beyond passive prompt-and-generate workflows to genuine human-AI collaboration requires solving open technical challenges in mutual agency, real-time latent steering, and dynamic evaluation.
Unconditional security against adaptive quantum adversaries marks a significant leap in certified randomness protocols, challenging previous assumptions and limitations.
Most AI fairness evaluations overlook critical uncertainty factors, risking misinterpretation of bias impacts in hiring systems.
Current MLLMs can only identify and explain environmental risks in secure authentication scenarios with limited effectiveness, revealing a significant vulnerability in their design.
Slasher can dynamically adjust datacenter power usage, ensuring operational resilience while safeguarding workload performance during critical power events.
Amortizing planning in latent world models leads to an order-of-magnitude reduction in planning time while boosting success rates across multiple benchmarks.
Iterative experience can boost LLM performance by over 5%, transforming how we evaluate and enhance model capabilities in real-time.
Cloning synthetic demonstrations alone caps performance, but combining it with reinforcement learning unlocks superior real-world manipulation capabilities.
A modular deep learning framework that enhances neuroimaging task performance by leveraging optimal task sequences from over 49,000 MRIs.
Relying on a single simulator in multi-agent RL leads to dangerous mode collapse, but innovative solutions can boost generalization and performance by up to 14%.
Flex-$\pi$ achieves unprecedented efficiency in bimanual manipulation by jointly learning from RGB and 3D geometry without extra training costs.
Positional bias in LLMs can drastically alter predictions, with accuracy and stability often at odds, challenging our understanding of model reliability in ordinal classification tasks.
The most effective playback similarity metric, CLEWS, is also the least expensive to implement, revolutionizing evaluation strategies in music transcription.
Factual hallucinations in LLM chess commentary are alarmingly high, with smaller models making over 40% incorrect claims, even after tool augmentation.
Uncertainty alone can mislead expert activation; VI-MoLE redefines routing by focusing on certified value-of-information, leading to more informed and efficient model decisions.
Cocktail-Talker enables spoken dialog systems to navigate the chaos of multi-speaker interactions, achieving selective engagement in noisy environments.
Current vision-language models fail to achieve embodied self-awareness, with none surpassing a 16.8% success rate in real-world interaction tasks.
OrganLens achieves unprecedented organ-specific representation learning in CT scans, significantly boosting diagnostic accuracy without requiring external segmentation.
Activating experts based on token uncertainty allows CARE to outperform traditional fixed routing methods while using fewer resources.
Offline policy-guided expert routing can significantly enhance biomedical image analysis, enabling robust performance without costly retraining.
MicroZoom synthesizes gigapixel-resolution images that maintain material authenticity and structural integrity, even from low-quality inputs.
Modular robot policies synthesized through natural language corrections outperform traditional black-box models, enabling interpretable and adaptable robotic behaviors.
Courts are relying on outdated legal doctrines to manage AI risks, leaving many harms unaddressed and creating a piecemeal governance landscape.
LLMs are overzealous tutors, intervening too soon and too often, which may undermine true learning and cognitive engagement.
Language models can master complex linguistic structures like unlike coordination without ever being explicitly trained on them, challenging traditional views on linguistic exposure.
Recovery routing can outperform escalation strategies by leveraging execution feedback, achieving a higher solve rate at only 35% of the typical recovery cost.
Achieving per-class coverage under distribution shifts is not only inefficient but also fundamentally impossible without sufficient target labels, revealing a hidden cost in label complexity.
TraversRL transforms pedestrian pathway generation by modeling it as a sequential decision-making process, yielding more reliable networks than traditional segmentation methods.
Morning and late-night routine volatility emerges as a critical digital biomarker for postpartum depression risk, revealing insights into maternal mental health dynamics.
Poisoning pretraining data through public discussion interfaces poses a significant threat, with the potential for undetectable harmful behaviors in language models.
Existing payloads in memory can compromise future agent behavior, revealing a critical vulnerability in memory-based AI systems.
Adaptive memory management can boost LLM task success by over 15 points while slashing token usage by up to 20%.
UrbanAgent outperforms traditional methods by leveraging multi-agent reasoning to tackle cross-modal inconsistencies in urban profiling tasks.
Pythia's autonomous prompt optimization outperforms traditional lexicon-based methods in clinical symptom detection, especially in low-prevalence scenarios.
Interaction scaling reveals a powerful new dimension of model performance that consistently outperforms traditional reasoning and sampling methods by leveraging real-time feedback.
Remote labs could redefine hands-on learning by ensuring students gain critical industry skills while studying from anywhere.
Remote access to real FPGA boards via a virtual breadboard could revolutionize hands-on learning in digital design courses.
Reasoning for control can be transformed into an adaptive, iterative process that leverages a latent memory structure, yielding superior performance in complex tasks.
Penalizing the decision-making path while rewarding the outcome can drastically reduce operational violations in real-world agent interactions.
MedPMC transforms millions of clinical articles into a goldmine of high-fidelity image-text pairs, dramatically improving model performance and clinical relevance.
High guidance in classifier-free guidance can destabilize models, but a simple adjustment can stabilize performance without extra computational cost.
RPAM reveals a robust connection between upstream evaluations and real-world biases, outperforming traditional downstream metrics in assessing language model associations.
State-of-the-art Vision-Language Models fall short in real-world robotic applications, revealing critical gaps in their reasoning capabilities.
Transforming attention heads into interpretable feature detectors could revolutionize how we understand and trust transformer models.
A unified decision process for multi-modal reasoning reveals that joint optimization of text and image generation can dramatically enhance performance in complex reasoning tasks.
Separable graphs reveal a hidden structure in graphical models that could unify diverse independence frameworks and streamline model identification.
A systematic comparison of software licenses reveals hidden attributes that could redefine how developers choose and enforce licensing terms.
A transformer with explicit fuzzy logic not only matches baseline performance but also reveals how it interprets grammatical structures, making model behavior legible.
Fairness evaluations of LLMs may be misleading, as models show a dramatic drop in ethical behavior when demographic cues are less explicit.
Claw-like agents are vulnerable to severe security breaches, with malicious plugins achieving a 100% success rate in attacks.
Human feedback can transform AI-generated notes into more helpful content, but these collaborative efforts still struggle for visibility compared to traditional notes.
Coding-agent workloads reveal surprising inefficiencies, with high cache hit rates but significant opportunities for optimization in LLM serving.
Agents can achieve a remarkable 19.4% harm rate while maintaining principal loyalty, but improving one aspect of their performance inevitably compromises another.
Touch is not just an add-on; it fundamentally enhances object representation, leading to dramatic improvements in physical property estimation and manipulation tasks.
Evolved agents can learn to dynamically coordinate multiple retrieval strategies, leading to a remarkable 19.6-point performance boost in multimodal document reasoning.
Retaining the original question in multilingual reasoning cascades can dramatically enhance performance, revealing that context is key to effective translation and reasoning.
MIRAGE successfully turns the tables on image editing systems by leveraging their own moderation processes to block unauthorized manipulations.
LLMs can generate millions in exploit profits, yet struggle to effectively patch vulnerabilities in smart contracts, revealing a critical gap in security capabilities.
Bridging the gap between informal and formal mathematics, TheoremGraph reveals 18.3 million dependencies that can enhance mathematical search and reasoning.
A learned continuous communication channel can dramatically enhance real-time game performance by bridging the gap between slow reasoning and fast reaction models.
Quantum pseudorandom states can only be stretched to a limited extent, revealing a stark contrast with classical counterparts.
Bad prompts can lead to a staggering 40% drop in LLM performance, revealing a critical vulnerability in in-context learning.
Sparse validation in AI-assisted interviews can dramatically improve data accuracy, revealing when and how to leverage structured questions effectively.
Linguistic analysis reveals that geographic distance between Tang poets correlates with distinct poetic styles, challenging assumptions about uniformity in literary expression. WHY_IT MATTERS: This research could transform our understanding of regional influences in historical literature and enhance the application of machine learning in humanities scholarship.
Training data diversity is the secret sauce that boosts agentic model performance, with OpenThoughts-Agent achieving a notable accuracy leap over existing benchmarks.
A fast, compact predictor that remains on the learned manifold for up to 1000 steps, significantly reducing long-horizon error compared to traditional methods.
Concordia achieves fault tolerance for LLM inference by seamlessly integrating persistent kernel checkpointing, enabling rapid recovery without CPU bottlenecks.
Tmax sets a new standard for terminal agent performance with a surprisingly simple RL recipe that outshines larger models.
Current embedding models misinterpret mathematical equivalence, grouping statements by terminology rather than content, but a new contrastive learning approach can bridge this gap.
Users adapt their AI detection strategies in real-time, reflecting both the evolution of generative models and shifting social dynamics.
Achieving 98.6% tracking efficiency with a 0.8% fake rate, HEPTv2 revolutionizes particle tracking by eliminating the need for graph construction and auxiliary processing.
Achieving a 7.6× speedup in distributed 3D scene reconstruction without sacrificing quality could redefine efficiency benchmarks in computer graphics.
EHR-based datasets may misrepresent suicidality by oversimplifying diverse clinical contexts into misleading labels.
Achieving state-of-the-art performance in 3D motion forecasting, MolmoMotion reveals that language instructions can significantly enhance trajectory predictions across various object categories and motion types.
IUU+DB transforms fragmented evidence of illegal fishing and labor abuses into actionable insights, revealing critical hotspots and trends that could reshape fisheries management.
LLMs can seamlessly generate formal proofs with a new system that mimics natural mathematical language, bridging the gap between human intuition and formal verification.
Coding agents can achieve up to 3.5x faster completion times with a new KVCache management strategy that understands their unique workload patterns.
Reprinted visual content in historic newspapers reveals intricate patterns of information exchange, showcasing the virality of images in American media history.
LLMs can autonomously optimize AI models for embedded devices, achieving 250x compression with minimal accuracy loss—something human experts struggle to match.
LLMs can lose over 30% of their accuracy in medical judgment when exposed to misleading contexts, revealing a critical vulnerability in their deployment for health advice.
Active learning can cut the data needed for accurate dynamics discovery by focusing on the most informative regions, outperforming random sampling by a significant margin.
M* achieves up to 2.9x lower real-time factor and 2.7x higher throughput for text-to-speech tasks, revolutionizing how we serve complex multimodal AI models.
By integrating frequent directions matrix sketching, EOFD-MLogB slashes the computational costs of multinomial logistic bandits without sacrificing performance.
Piper enables seamless integration of cutting-edge parallelism strategies, achieving performance parity with existing methods while unlocking new efficiency gains.
Optimized fuel formulations can significantly reduce pollutant emissions while meeting stringent aviation standards, outperforming existing blends across key metrics.
Injecting diversity at the specification level can significantly enhance output variety while preserving quality, challenging conventional methods that often fall short.
Reusing training data during inference can boost imitation learning performance by up to 46%, reshaping how we approach generalization in AI systems.
Explicitly inferring mental states under uncertainty leads to more thoughtful dialogue than unrestricted information access, challenging conventional wisdom in AI interactions.
Torque Adaptation Module enables zero-shot robust manipulation across different robots without the need for extensive retraining or domain randomization.
Success in long-horizon tasks hinges more on an agent's iterative persistence than on the quality of its initial solution.
Imaginative Perception Tokens boost spatial reasoning in VLMs, achieving a 3.4% accuracy gain on Multiview Counting while outperforming traditional training methods.
Achieving a 3% boost in identity tracking accuracy in thermal video by focusing on trajectory relinking rather than complex models could redefine best practices in MOT.
Achieving a 4.2x latency reduction in long-form ASR without sacrificing accuracy could revolutionize real-time speech applications.
Steering imaginations in video world models can reveal critical failure points in robotic actions that traditional methods might overlook.
Recurrent memory can be added to transformers at scale with minimal parameter overhead and no performance penalty by reusing existing hidden states and training with interleaved parallel updates.
The best LLM to answer a question isn't always the best LLM to *teach* the answer, and matching the "difficulty" of the explanation to the student's current abilities yields better learning.
LLMs can resolve merge conflicts nearly as well as Google's best, but still fail in over 40% of cases, revealing a surprising bottleneck in automating software development.
Stop banning GenAI in STEM assessments: this framework shows you how to thoughtfully integrate it to actually *improve* learning outcomes.