Search papers, labs, and topics across Lattice.

Top-tier US AI research university. Strong in NLP, ML systems, and computer vision.
100
0
0
LLMs are overzealous tutors, intervening too soon and too often, which may undermine true learning and cognitive engagement.
Language models can master complex linguistic structures like unlike coordination without ever being explicitly trained on them, challenging traditional views on linguistic exposure.
Recovery routing can outperform escalation strategies by leveraging execution feedback, achieving a higher solve rate at only 35% of the typical recovery cost.
TraversRL transforms pedestrian pathway generation by modeling it as a sequential decision-making process, yielding more reliable networks than traditional segmentation methods.
Achieving per-class coverage under distribution shifts is not only inefficient but also fundamentally impossible without sufficient target labels, revealing a hidden cost in label complexity.
Morning and late-night routine volatility emerges as a critical digital biomarker for postpartum depression risk, revealing insights into maternal mental health dynamics.
Poisoning pretraining data through public discussion interfaces poses a significant threat, with the potential for undetectable harmful behaviors in language models.
Existing payloads in memory can compromise future agent behavior, revealing a critical vulnerability in memory-based AI systems.
UrbanAgent outperforms traditional methods by leveraging multi-agent reasoning to tackle cross-modal inconsistencies in urban profiling tasks.
Adaptive memory management can boost LLM task success by over 15 points while slashing token usage by up to 20%.
Pythia's autonomous prompt optimization outperforms traditional lexicon-based methods in clinical symptom detection, especially in low-prevalence scenarios.
Interaction scaling reveals a powerful new dimension of model performance that consistently outperforms traditional reasoning and sampling methods by leveraging real-time feedback.
Remote labs could redefine hands-on learning by ensuring students gain critical industry skills while studying from anywhere.
Remote access to real FPGA boards via a virtual breadboard could revolutionize hands-on learning in digital design courses.
Reasoning for control can be transformed into an adaptive, iterative process that leverages a latent memory structure, yielding superior performance in complex tasks.
MedPMC transforms millions of clinical articles into a goldmine of high-fidelity image-text pairs, dramatically improving model performance and clinical relevance.
Penalizing the decision-making path while rewarding the outcome can drastically reduce operational violations in real-world agent interactions.
High guidance in classifier-free guidance can destabilize models, but a simple adjustment can stabilize performance without extra computational cost.
State-of-the-art Vision-Language Models fall short in real-world robotic applications, revealing critical gaps in their reasoning capabilities.
RPAM reveals a robust connection between upstream evaluations and real-world biases, outperforming traditional downstream metrics in assessing language model associations.
Transforming attention heads into interpretable feature detectors could revolutionize how we understand and trust transformer models.
A unified decision process for multi-modal reasoning reveals that joint optimization of text and image generation can dramatically enhance performance in complex reasoning tasks.
Separable graphs reveal a hidden structure in graphical models that could unify diverse independence frameworks and streamline model identification.
A transformer with explicit fuzzy logic not only matches baseline performance but also reveals how it interprets grammatical structures, making model behavior legible.
A systematic comparison of software licenses reveals hidden attributes that could redefine how developers choose and enforce licensing terms.
Fairness evaluations of LLMs may be misleading, as models show a dramatic drop in ethical behavior when demographic cues are less explicit.
Agents can achieve a remarkable 19.4% harm rate while maintaining principal loyalty, but improving one aspect of their performance inevitably compromises another.
Coding-agent workloads reveal surprising inefficiencies, with high cache hit rates but significant opportunities for optimization in LLM serving.
Human feedback can transform AI-generated notes into more helpful content, but these collaborative efforts still struggle for visibility compared to traditional notes.
Claw-like agents are vulnerable to severe security breaches, with malicious plugins achieving a 100% success rate in attacks.
Evolved agents can learn to dynamically coordinate multiple retrieval strategies, leading to a remarkable 19.6-point performance boost in multimodal document reasoning.
Touch is not just an add-on; it fundamentally enhances object representation, leading to dramatic improvements in physical property estimation and manipulation tasks.
Retaining the original question in multilingual reasoning cascades can dramatically enhance performance, revealing that context is key to effective translation and reasoning.
Bridging the gap between informal and formal mathematics, TheoremGraph reveals 18.3 million dependencies that can enhance mathematical search and reasoning.
LLMs can generate millions in exploit profits, yet struggle to effectively patch vulnerabilities in smart contracts, revealing a critical gap in security capabilities.
MIRAGE successfully turns the tables on image editing systems by leveraging their own moderation processes to block unauthorized manipulations.
Sparse validation in AI-assisted interviews can dramatically improve data accuracy, revealing when and how to leverage structured questions effectively.
Linguistic analysis reveals that geographic distance between Tang poets correlates with distinct poetic styles, challenging assumptions about uniformity in literary expression. WHY_IT MATTERS: This research could transform our understanding of regional influences in historical literature and enhance the application of machine learning in humanities scholarship.
Bad prompts can lead to a staggering 40% drop in LLM performance, revealing a critical vulnerability in in-context learning.
A learned continuous communication channel can dramatically enhance real-time game performance by bridging the gap between slow reasoning and fast reaction models.
Training data diversity is the secret sauce that boosts agentic model performance, with OpenThoughts-Agent achieving a notable accuracy leap over existing benchmarks.
Quantum pseudorandom states can only be stretched to a limited extent, revealing a stark contrast with classical counterparts.
A fast, compact predictor that remains on the learned manifold for up to 1000 steps, significantly reducing long-horizon error compared to traditional methods.
Concordia achieves fault tolerance for LLM inference by seamlessly integrating persistent kernel checkpointing, enabling rapid recovery without CPU bottlenecks.
Tmax sets a new standard for terminal agent performance with a surprisingly simple RL recipe that outshines larger models.
Current embedding models misinterpret mathematical equivalence, grouping statements by terminology rather than content, but a new contrastive learning approach can bridge this gap.
Users adapt their AI detection strategies in real-time, reflecting both the evolution of generative models and shifting social dynamics.
Achieving 98.6% tracking efficiency with a 0.8% fake rate, HEPTv2 revolutionizes particle tracking by eliminating the need for graph construction and auxiliary processing.
Achieving a 7.6× speedup in distributed 3D scene reconstruction without sacrificing quality could redefine efficiency benchmarks in computer graphics.
EHR-based datasets may misrepresent suicidality by oversimplifying diverse clinical contexts into misleading labels.
Achieving state-of-the-art performance in 3D motion forecasting, MolmoMotion reveals that language instructions can significantly enhance trajectory predictions across various object categories and motion types.
IUU+DB transforms fragmented evidence of illegal fishing and labor abuses into actionable insights, revealing critical hotspots and trends that could reshape fisheries management.
LLMs can seamlessly generate formal proofs with a new system that mimics natural mathematical language, bridging the gap between human intuition and formal verification.
Coding agents can achieve up to 3.5x faster completion times with a new KVCache management strategy that understands their unique workload patterns.
Reprinted visual content in historic newspapers reveals intricate patterns of information exchange, showcasing the virality of images in American media history.
LLMs can autonomously optimize AI models for embedded devices, achieving 250x compression with minimal accuracy loss—something human experts struggle to match.
LLMs can lose over 30% of their accuracy in medical judgment when exposed to misleading contexts, revealing a critical vulnerability in their deployment for health advice.
By integrating frequent directions matrix sketching, EOFD-MLogB slashes the computational costs of multinomial logistic bandits without sacrificing performance.
Active learning can cut the data needed for accurate dynamics discovery by focusing on the most informative regions, outperforming random sampling by a significant margin.
M* achieves up to 2.9x lower real-time factor and 2.7x higher throughput for text-to-speech tasks, revolutionizing how we serve complex multimodal AI models.
Injecting diversity at the specification level can significantly enhance output variety while preserving quality, challenging conventional methods that often fall short.
Piper enables seamless integration of cutting-edge parallelism strategies, achieving performance parity with existing methods while unlocking new efficiency gains.
Optimized fuel formulations can significantly reduce pollutant emissions while meeting stringent aviation standards, outperforming existing blends across key metrics.
Reusing training data during inference can boost imitation learning performance by up to 46%, reshaping how we approach generalization in AI systems.
Torque Adaptation Module enables zero-shot robust manipulation across different robots without the need for extensive retraining or domain randomization.
Explicitly inferring mental states under uncertainty leads to more thoughtful dialogue than unrestricted information access, challenging conventional wisdom in AI interactions.
Success in long-horizon tasks hinges more on an agent's iterative persistence than on the quality of its initial solution.
Imaginative Perception Tokens boost spatial reasoning in VLMs, achieving a 3.4% accuracy gain on Multiview Counting while outperforming traditional training methods.
Achieving a 3% boost in identity tracking accuracy in thermal video by focusing on trajectory relinking rather than complex models could redefine best practices in MOT.
Achieving a 4.2x latency reduction in long-form ASR without sacrificing accuracy could revolutionize real-time speech applications.
Steering imaginations in video world models can reveal critical failure points in robotic actions that traditional methods might overlook.
Recurrent memory can be added to transformers at scale with minimal parameter overhead and no performance penalty by reusing existing hidden states and training with interleaved parallel updates.
The best LLM to answer a question isn't always the best LLM to *teach* the answer, and matching the "difficulty" of the explanation to the student's current abilities yields better learning.
LLMs can resolve merge conflicts nearly as well as Google's best, but still fail in over 40% of cases, revealing a surprising bottleneck in automating software development.
Stop banning GenAI in STEM assessments: this framework shows you how to thoughtfully integrate it to actually *improve* learning outcomes.
Unlock the long tail of autonomous driving scenarios: Sensor2Sensor turns readily available dashcam footage into high-fidelity, multi-modal sensor data, bridging the gap between data scarcity and the need for robust AV training.
Ophthalmic VQA models can be made more accurate and transparent by explicitly grounding them in spatially-localized lesion evidence, a crucial step towards clinical interpretability.
Training a foundation model on a trillion minutes of wearable sensor data unlocks surprisingly accurate predictions across a wide range of health conditions, even with limited labeled data.
Forget Gaussian noise - modeling the *decay* of user interest with a custom "burn-down" diffusion process unlocks better personalized recommendations.
Distributional regret bounds, which quantify the probability of exceeding different regret levels, are now achievable with a UCBVI-style algorithm, confirming a long-standing conjecture for multi-armed bandits.
AI data annotation companies are publicly framing human expertise as a commodity ripe for disruption, potentially devaluing traditional forms of knowledge and institutional authority.
LLMs can't rebuild software from scratch, even for widely used programs like FFmpeg and SQLite, revealing a critical gap in their ability to make high-level software architecture decisions.
Open-sourcing a VLA model that beats closed-source giants on embodied reasoning tasks could finally make real-world robot deployment practical.
Unlock collaborative AI development in genomics without compromising patient privacy: this framework lets multiple institutions jointly train synthetic data generators on sensitive RNA-seq data using MPC and DP.
Traditional research papers are costing AI agents reproducibility and understanding, but a new "Agent-Native" format that captures the full messy research process boosts performance by up to 20%.
Multimodal models can now achieve state-of-the-art performance in real-world tasks like document understanding and audio-video comprehension with significantly reduced inference latency thanks to novel token-reduction techniques.
LLM agents struggle to maintain performance in multi-day collaborative tasks, dropping significantly after just one environmental update, revealing a critical gap in adaptation to evolving real-world conditions.
The fragmented field of world modeling can now be unified under a "levels x laws" taxonomy, revealing critical gaps in autonomous model revision and decision-centric evaluation.
A surprisingly simple tweak to Hartigan's k-means algorithm unlocks another 2-5% accuracy boost, especially when clustering high-dimensional data.
Modular training with BAR allows independent updates of domain experts, achieving superior performance without the pitfalls of catastrophic forgetting.
Kernel launch overhead is a bigger bottleneck than you think: GPUOS achieves up to 15.3x speedup by fusing operations at runtime.
RosettaSearch recovers up to 68% more structural fidelity in protein designs, transforming how we optimize sequences beyond traditional single-pass methods.
Geometric matrix interpolation reveals hidden common structures in multi-view data, offering a new lens for multi-manifold learning.
Massively multilingual NER just got easier: UNER v2 offers a standardized benchmark for evaluating LLMs across diverse languages.
LLMs are twice as likely as humans to repeat the same support tactic in a conversation, but a simple RL reward for tactic novelty can fix it.
Achieving robust brain decoding across subjects without any retraining could revolutionize how we interpret neural signals in diverse populations.
Forget training on closed sets: WildDet3D leverages geometric cues and diverse prompts to achieve SOTA 3D object detection across 13.5K categories in the wild.
MLLMs can be tricked into missing 90% of harmful content simply by encoding it in images that humans can easily read.
Get 80% of your prompt length back without sacrificing accuracy using a diffusion-based pruning method that can mask multiple tokens at once.
Serving both image and video diffusion models on the same hardware? GENSERVE's step-level preemption and dynamic resource allocation can boost your service level agreement (SLA) attainment by up to 44%.