Search papers, labs, and topics across Lattice.

MIT's Computer Science and Artificial Intelligence Laboratory. One of the largest and oldest AI labs in academia.
100
1
0
Choosing AI for emotional support not only enhances immediate satisfaction but also reshapes long-term preferences away from human interaction.
Even with an impressive $R^2$ of 0.75, EO-ML methods can mislead policymakers due to inherent uncertainties, underscoring the need for robust uncertainty quantification in poverty mapping.
Achieving a 40% reduction in character error rate, this syllable-level UASR framework unlocks new potential for low-resource language recognition without costly phoneme resources.
Achieving superior semantic segmentation performance, Contextrast++ tackles long-tailed distribution issues and enhances context awareness without additional inference overhead.
Selective re-scanning in recurrent networks can drastically reduce memory usage while improving task performance, challenging the conventional wisdom of fixed-size state fidelity.
A universal quadratic model reveals that diverse neural architectures share a common training dynamic, leading to predictable power law behaviors in learning.
SoftWater slashes quantization error by up to 8.3x while maintaining near-lossless performance in language models, revolutionizing softmax layer efficiency.
FiGuRO reveals that effective intrinsic dimension estimation can emerge as a byproduct of optimizing low-rank projections, transforming how we approach multi-modal representation learning.
AMIE (Video) outperformed human physicians in critical clinical tasks, signaling a leap toward AI that can effectively engage in complex medical consultations.
SVI-DAG outperforms existing Bayesian methods by effectively quantifying uncertainty in causal inference while leveraging prior knowledge and edge dependencies.
Spectral fingerprints can differentiate between unique molecular structures with identical 2D connectivity, revolutionizing how we assess chemical similarity.
Social influence can lead clinical decision support agents to adopt incorrect answers at alarming rates, revealing a critical flaw in multi-agent oversight.
Dropping echocardiogram data nearly doubles error rates in cardiac AI models, revealing critical modality dependencies that could impact clinical decisions.
Language model agents struggle with oncall RCA, achieving only 25.3% accuracy on realistic tasks, revealing a critical readiness gap for production environments.
Per-group threshold optimization consistently outperforms traditional calibration methods, revealing that variance, not averages, is crucial for accurate fairness auditing in clinical risk models.
SciFigAlign achieves a remarkable 59% reduction in error over traditional LLM-based scoring methods by grounding figure assessments in manuscript context.
Energy estimates for LLM inference can be achieved without direct measurements, revealing insights into the environmental impact of AI systems.
Proprietary MLLMs may achieve high diagnostic accuracy, but they still struggle with reliable clinical reasoning, revealing significant gaps in their practical utility.
Modular robot policies synthesized through natural language corrections outperform traditional black-box models, enabling interpretable and adaptable robotic behaviors.
Ms.Forcing achieves a 39.6% speedup in streaming video generation while significantly enhancing quality by intelligently adapting spatial granularity to noise levels.
PerfAgent doubles the rate of expert-level code optimizations by leveraging profiler-guided feedback, revealing hidden performance bottlenecks that traditional methods miss.
Remarkably, even incomplete or ill-conditioned descriptors can yield accurate atomic reconstructions, challenging the conventional wisdom about the necessity of high-dimensional feature sets.
ATLAS achieves over 500-fold efficiency in sampling amorphous materials while maintaining less than 0.2% free energy error, revolutionizing the approach to material design.
Achieving optimal convergence rates for nonconvex optimization by transforming it into a static regret minimization problem could revolutionize how we design adaptive optimizers.
Verifying audio authenticity through a dual alignment likelihood ratio test reveals a faster and more robust alternative to traditional tampering detection methods.
AutoSynthesis achieves expert-level meta-analysis with automated precision, making evidence synthesis scalable and accessible.
DriftWorld achieves 17x faster rollouts than diffusion models, revolutionizing real-time planning for robotic manipulation.
ATLAS transforms LLMs from unreliable code generators to trusted partners in analog design, successfully producing SAR ADCs that meet rigorous simulation standards.
Skip connections and normalization layers are not just about controlling magnitude; they are crucial for preserving gradient rank and influencing model performance.
Data synergy can either amplify or diminish model performance, revealing that the right dataset combinations are crucial for optimal language model training.
Hard interventions can reveal causal structures even when traditional assumptions of faithfulness fail, challenging the status quo in causal discovery.
Multimodal unlearning could revolutionize how we handle sensitive data in AI, enabling targeted removal without sacrificing model performance.
Silent policy violations in tool-using LLMs can be mitigated by deterministic gates, improving success rates by over 12 percentage points in critical tasks.
Frequency usage in transformers is not random; it’s intricately tied to the data’s dependency structure, revealing a data-driven mechanism behind RoPE's emergent behavior.
Deployment rules can shift multi-agent AI outcomes dramatically, with fatality rates varying by up to 58 percentage points based solely on the chosen rule.
LLMs can be trained to negotiate like expert agents, extracting significantly higher surpluses by strategically exploring buyer markets rather than fixating on immediate bids.
The dominance of a few countries and institutions in AI bias research risks creating a narrow lens through which fairness is defined and addressed.
Nearly 80% of AI-generated pull requests are submitted concurrently, raising critical questions about collaboration efficiency and merge conflicts in AI coding agents.
LLM-driven program synthesis can automate EEG feature engineering while ensuring interpretability and high detection accuracy.
Projected reads in PatchOptic not only cut token costs but also ensure that local updates remain valid in the context of shared-state workflows.
Low diversity in training data can lead to substantial performance drops in language models, revealing a critical oversight in data augmentation practices.
The choice of performance metrics could determine whether AI capabilities remain concentrated among the wealthy or proliferate across a broader developer base.
NNPs can achieve near-chemical accuracy in enzyme catalysis predictions with less than 1,000 system-specific data points, revolutionizing the efficiency of mechanistic studies.
DiscoPER not only automates hypothesis generation but also self-analyzes its discoveries, revealing hidden patterns and expanding the search space in unprecedented ways.
Adversarial training with human demonstrations can significantly enhance the quality and diversity of language model outputs while preserving accuracy.
SLIM-RL achieves state-of-the-art performance on math and code tasks with nearly half the training samples required by traditional trajectory-aware methods.
Automating freeway network extraction from OSM can cut analyst effort by two-thirds, making large-scale freeway simulations feasible.
Fixed counterfactual explanations can lead LMs to generate more accurate introspections about their behaviors, even as those behaviors change over time.
Historical failure records can be transformed into diverse testing scenarios for autonomous driving, revealing critical system vulnerabilities with minimal effort.
Relying on text-based rationales for dementia classification can actually degrade performance, revealing the need for more sophisticated approaches like DeTAiL.
Evaluating creative AI requires recognizing that professional disagreement reflects genuine taste differences, not just noise in measurement.
Generative AI agents can reveal how personalization algorithms amplify toxic content in ways that vary dramatically by user ideology.
A unified platform for molecular machine learning that supports 100 elements and incorporates uncertainty quantification could democratize access to advanced chemical property predictions.
LLMs reveal a modular cognitive architecture strikingly similar to that of the human brain, challenging our understanding of intelligence across different systems.
Memory consistency in video generation models falters significantly when objects disappear, with state-of-the-art models struggling to recover updated states upon reappearance.
ArBG achieves a remarkable 60% reduction in zero-shot energy error for peptide systems, challenging the dominance of flow-based sampling methods.
Dynamic frame rates in audio autoencoders can drastically enhance efficiency, allowing for smarter resource allocation in neural compression tasks.
FracEvent achieves superior event timing and downstream performance by accurately modeling pixel dynamics, outperforming traditional simulators.
GPUSparse achieves a staggering 235x speedup over traditional CPU methods while maintaining exact scoring, revolutionizing real-time retrieval efficiency.
Complex manipulation capabilities can be achieved by dynamically composing simple behaviors, leading to unprecedented precision and adaptability in real-world tasks.
Achieving a staggering 220x speedup in MaxSim scoring while preserving exact retrieval quality could revolutionize the efficiency of multi-vector retrieval systems.
Fixed exponents in neural scaling laws reveal that optimizing coefficients could unlock significant performance gains in large language models.
Current evaluation metrics for AI-powered AAC systems overlook the intersectional nuances of user needs, risking ineffective communication solutions.
Cross-lingual exploration can unlock hidden knowledge in LLMs, improving factual recall and consistency across 17 languages.
Gaussian soft labels boost landmark detection performance by 7% over traditional hard labels, revealing the importance of modeling annotation variability.
Decentralized traffic management for autonomous aircraft can achieve high performance without centralized coordination, adapting seamlessly to complex environments.
Turn-final words are not just longer; they provide a crucial prosodic cue for predicting conversational turns, localized mainly in the final syllable.
Over 200,000 synthesized words reveal the intricate relationship between articulatory gestures and acoustic landmarks in speech.
Trajectory mining reveals skill structures but fails to translate these insights into meaningful performance gains for downstream policies.
Balancing productivity and stability reveals that stronger synchronization can paradoxically increase systemic fragility in multi-agent systems.
Machine learning can transform 2DES by extracting maximum insights from limited data while guiding experimental design for improved accuracy.
WalkOCC achieves superior sidewalk occupancy prediction by leveraging unpaired monocular images, eliminating the need for costly 3D annotations.
Training climate emulators on a single optimized scenario can outperform those trained on six standard pathways, challenging the notion that more data always leads to better performance.
Turing-RL reveals that training user simulators for indistinguishability can dramatically improve their performance in simulating human interactions.
Executable programs can now replace attention heads in transformers with minimal performance loss, achieving over 75% similarity to original patterns.
Nonuniform width allocation in transformers can lead to a 22% reduction in FLOPs while enhancing language modeling performance.
The largest-ever verification campaign for Rust's standard library reveals significant vulnerabilities in unsafe code, underscoring the need for robust static verification methods.
MixTIME reveals that integrating diverse image modalities can significantly enhance the precision of immune biomarker predictions in oncology, outperforming traditional single-modality approaches.
Retaining visual figures in skill artifacts boosts CUA performance by over 23 points, proving that seeing is believing in agent training.
Post-congestion pricing, NYC saw a surge in transit ridership while overall travel demand fell, highlighting the complex interplay of urban mobility and pricing policies.
Dynestyx transforms the landscape of Bayesian analysis by making advanced state-space modeling techniques accessible to practitioners through a unified interface.
LLMs show significant variability in the actionability of their UX critiques, with some models outperforming others across different product categories.
Coordination-free DIDs could revolutionize decentralized identity management by enabling instantaneous updates without the overhead of global ordering.
Naively scaling context length in imitation learning is surprisingly robust, challenging previous assumptions about brittleness in policy performance.
LLMs can now interpret quantum operations, achieving state-of-the-art results in circuit synthesis while allowing natural language constraints.
Even state-of-the-art multimodal models struggle with reliability in clinical tool use, revealing critical gaps in AI agent performance.
The "curse of precision" reveals how reliance on AI-generated content can degrade model performance by homogenizing training data.
Current AI agents excel in structured tasks but falter at generating novel insights and tackling open-ended scientific challenges.
Action-chunking policies can lead to premature robot assistance, but a novel steering method effectively mitigates this issue, enhancing collaboration efficiency.
Optimizing input configurations can boost LLM performance in pathology tasks, closing the gap with specialized models and challenging assumptions about domain-specific training.
Standard stereo methods can produce 3D models of Martian terrain, but achieving reliable reconstruction demands careful consideration of domain-specific challenges.
State inertia in full-duplex spoken language models can lead to missed user input, but activation steering effectively mitigates this issue, boosting comprehension rates significantly.
Gradient inference can revolutionize how we tackle gradient estimation in complex probabilistic programs, enabling new state-of-the-art estimators that outperform traditional methods.
Systematic gaps in AI evaluation reporting are exposed, revealing inconsistencies that hinder reliable comparisons across thousands of models and benchmarks.
Despite promising engagement benefits, foundation model-based care robots struggle with reliability and lack robust evidence for clinical impact.
AI systems currently miss critical temporal and interpretive elements of clinical reasoning, limiting their effectiveness in real-world healthcare settings.
Current clinical AI systems often neglect the temporal dimension of patient care, limiting their effectiveness in longitudinal reasoning.
Design choices in agent memory systems can significantly shift operational costs, revealing critical trade-offs that impact long-horizon task performance.
Active exploration can dramatically enhance adults' ability to reason about complex causal relationships, but even with this advantage, they still struggle compared to simpler tasks.
Monte Carlo methods can now compute Steklov spectra orders of magnitude faster while handling complex, disconnected geometries in real-world datasets.