Search papers, labs, and topics across Lattice.
Carnegie Mellon's Machine Learning Department. Home to foundational work in statistical ML, deep learning, and robotics.
100
2
0
Free-form language reasoning transforms VLMs into powerful robotic reasoners, significantly boosting performance in complex manipulation tasks.
HSR boosts robot manipulation success rates by over 21% by leveraging hierarchical skill retrieval, even with minimal task-specific data.
GCA reduces communication overhead in federated learning by up to 99.15% while simultaneously enhancing data protection and improving model accuracy.
By making environment design a learnable process, SPADE unlocks a new frontier in self-improvement for language agents, leading to substantial performance gains across diverse tasks.
A 5-step sampling schedule can deliver nearly the same quality as a 50-step one, slashing inference costs by 90%.
Achieving a new upper bound of $\omega < 2.371177 could redefine our understanding of matrix multiplication efficiency.
Open-vocabulary monocular 3D detectors mislabel correctly localized objects due to prompt sensitivity, revealing a critical gap in their semantic understanding.
HarnessEval-W transforms world model evaluation from mere scoring to a transparent reasoning process that mirrors human judgment.
Closing the loop on robot manipulation failures, VLCP rewrites control code in real-time, achieving a tenfold increase in success rates over traditional methods.
Current AI-generated videos that mislead viewers are also the most challenging for existing detection systems to identify, revealing a critical vulnerability in misinformation defenses.
Explicitly optimizing for target function smoothness can dramatically enhance the performance of models on tabular data.
Social Gym reveals that even top-performing LLMs struggle with consistent social reasoning across diverse multi-agent scenarios.
A 'severity-first' approach to risk assessment in the EU AI Act could expose vulnerabilities to manipulation by AI providers, risking regulatory effectiveness.
High speech overlap isn't the primary challenge in cocktail-party scenarios; innovative audio-visual strategies and large language models can cut recognition errors by 57%.
SJRL not only overcomes collision challenges in multi-agent navigation but also adapts dynamically to real-world constraints, outperforming traditional methods in complex environments.
Skill-switching accuracy in LLMs drops significantly on complex tasks, but a new training approach boosts performance from 34.4% to 68.4% on challenging benchmarks.
Formal verification of a compiler for asynchronous dataflow could redefine the reliability and efficiency of parallel computing architectures.
Path-level pretraining in MultiPathFormer leads to a dramatic 59% improvement in wireless propagation estimations, reshaping the landscape of wireless foundation models.
The rise of AI in software development is rendering traditional software measurement assumptions obsolete, necessitating a radical rethink of how we validate our findings.
H2S achieves a remarkable 48.54 mAP in audio-visual instance segmentation, setting a new benchmark in the field.
No randomized polynomial-time algorithm can overcome the condition-number barrier in sparse least-squares optimization, confirming a long-standing conjecture.
Retaining the right contextual information can boost long-context modeling performance by over 20% compared to traditional methods.
Agentic Commerce World reveals that process-level evidence is crucial for accurately evaluating AI agents in dynamic market environments, challenging traditional reliance on final outcomes alone.
Code writing outperforms all other active learning methods in programming education, leading to significantly better learning outcomes.
Achieving high-quality reranking with a 30B MoE model is now feasible on an academic budget, outperforming traditional dense models in efficiency.
SRA can double the operational capacity of automated warehouses while slashing optimization time from hours to minutes.
GenAI could either reinforce linguistic hierarchies or become a powerful tool for promoting diversity in academic writing, depending on its design and governance.
Coding agents may boost productivity, but they risk diminishing developers' understanding and long-term coding skills.
VLMs can fabricate highly biased medical diagnoses based solely on demographic descriptors, revealing alarming implications for clinical trustworthiness.
Achieving over 99% fulfillment of observation requests in decentralized satellite scheduling reveals a transformative approach to large-scale DCOPs that outshines traditional methods.
Downstream performance hinges on SSL pre-training language, not the NAC training language, enabling efficient cross-language model reuse.
Proprietary MLLMs may achieve high diagnostic accuracy, but they still struggle with reliable clinical reasoning, revealing significant gaps in their practical utility.
Retrieval harm can significantly impact performance, and simple baselines often outperform complex models in Active RAG systems.
Most developers edit AI-generated code within 15 minutes, often discarding the original completions entirely, highlighting a critical gap in LLM training data.
Achieving 19 times lower inference latency than diffusion models, UniGen-AR redefines the efficiency of unified visual generation.
Adapters may store less than half the information expected, with their capacity heavily influenced by parameter placement rather than quantity.
Different views of the same problem reveal hidden reasoning paths, enabling VLMs to achieve unprecedented accuracy in multimodal reasoning tasks.
A lightweight vision-language model can effectively identify UI principle violations, achieving an impressive F1 score of 84% across multiple quality dimensions.
Attention-supervised finetuning reveals that model architecture significantly impacts the plausibility of explanations in media bias detection, challenging assumptions about scale and performance.
Token generation timing can leak critical architectural and optimization details from language models, exposing vulnerabilities in their deployment.
AniGS transforms static 3D reconstructions into dynamic, immersive environments by seamlessly integrating ambient motion without compromising structural fidelity.
Smoothing the loss landscape with chain-of-thought reasoning can significantly enhance LLMs' ability to navigate complex moral dilemmas.
A unified framework reveals that robust downgrading mechanisms can coexist with full-strength non-interference, transforming our approach to program security.
渭STM achieves high performance without sacrificing safety or usability, allowing for general data types and deferred aborts in transactions.
The PanAf-SBR dataset reveals that fine-grained social behaviour recognition in wild great apes can be significantly improved through targeted cross-dataset pre-training.
EvolvingWorld reveals that an open-schema framework can drastically enhance the coherence and depth of character and world interactions in long-horizon literary simulations.
Nine out of ten AI-selected modeling changes in materials science remain effective when tested on unseen data, showcasing the potential for reusable AI-driven discoveries.
Visualizing navigation structures transformed practitioners' approach to accessibility, shifting their mindset from compliance to design innovation.
LLMs struggle with historical analogy retrieval due to a focus on surface features, but the new CANA framework boosts their performance by 10% through causal understanding.
Missing treatment data can lead to biased policy estimates, but leveraging partially-observed units with MAR estimators offers a more efficient and valid approach.
EMAGN reduces the complexity of traffic forecasting models from quadratic to linear, enabling larger configurations without sacrificing performance.
Requential coding can compress billion-parameter models to sizes orders of magnitude smaller than traditional methods, revealing the hidden structure in datasets.
Indirect data poisoning can enable scientific fraud at an unprecedented scale, with a staggering 49.56% success rate in corrupting AI-driven research outputs.
TALRanker autonomously navigates the trade-off between efficiency and accuracy in document reranking, achieving state-of-the-art results without the latency of excessive tool calls.
Phonological features can be extracted from self-supervised speech models in under a minute, achieving state-of-the-art performance in phone segmentation and recognition.
LightCrafter achieves superior video relighting by integrating PBR with diffusion models, enabling intricate lighting control and long-form temporal consistency without the need for extensive training data.
Agents can now work independently on data changes while humans maintain oversight, revolutionizing collaborative data management.
Every program in the new Calf framework must preserve both abstraction and potential, revolutionizing how we approach cost verification in type theory.
Current LLMs only achieve 27.3% accuracy in reasoning about scientific lineage, revealing a critical gap in their compositional capabilities.
Multimodal unlearning could revolutionize how we handle sensitive data in AI, enabling targeted removal without sacrificing model performance.
Achieving up to 98.75% accuracy in detecting autism-related behaviors highlights the potential of sequence-based models over traditional CNNs in data-scarce environments.
Achieving trajectory-level differential privacy in adaptive streaming contexts without sacrificing performance is now feasible through an auditable buffering-aggregation approach.
Operational reframing emerges as a critical risk signal, revealing that compliance can vary significantly across models and scenarios, challenging the notion of stable safety metrics in multi-agent LLMs.
ExplAIner can express a diverse array of explanation types while ensuring efficient evaluation, transforming how we approach interpretability in machine learning models.
RABBiT can accurately predict brain responses to speech with just 10 minutes of participant data, outperforming traditional models and enabling scalable population-level studies.
Prompting decoder-only models outperforms fine-tuning methods, achieving unprecedented effectiveness in ranking case-law sentences for statutory term retrieval.
GPT-4o can identify antisemitic incidents but requires better prompts to enhance its classification accuracy.
Many robotic policies that seem successful in manipulation tasks actually compromise safety, with SoftVTBench revealing a stark contrast between goal completion and physical safety metrics.
Static prompts in RL training can hinder performance, but LLM-as-a-Tutor dynamically adapts them to match policy capabilities, leading to superior outcomes.
None of the 30 LLM agents evaluated in CausalGame demonstrated reliable causal thinking, revealing a critical gap in AI's ability to perform scientific reasoning.
PACE-Bench predicts agentic performance with remarkable accuracy while slashing evaluation costs to a fraction of traditional methods.
Better attribution in generative music could significantly boost creator welfare and reshape platform compensation strategies.
Fixed-point flows enable a leap in performance for language models, outperforming state-of-the-art methods in one- and few-step generation tasks.
MoVA achieves superior video-text alignment by disentangling evolving visual concepts from static textual descriptions, outperforming existing models in handling long sequences.
Direct thread communication in MPLMs cuts context requirements by half, revolutionizing how LLMs tackle complex reasoning tasks.
Robots can now autonomously learn from their failures, boosting success rates by over 17% without human intervention.
Traditional metrics fail to capture the true memory capabilities of LLMs, exposing a critical gap in how we assess their deployment readiness.
FARS challenges the boundaries of automated research by producing 166 papers across 67 topics, revealing both its potential and pitfalls in AI-driven science.
Allowing language models to explore unsafe reasoning can actually enhance their ability to discern harmful from harmless prompts, reducing over-refusal without sacrificing safety.
ELASTIC reduces wall-clock latency by 34% while matching the success rates of the best-performing models in real-world robot manipulation.
Even top-performing AI models struggle with PowerPoint tasks, achieving only 45% success rates despite a robust evaluation framework that rewards nuanced performance.
Personalized fine-tuning of ASR models can reduce word error rates for dysarthric speech to as low as 9.7%, transforming communication for affected individuals.
Steering vectors can transform how we control language models, paving the way for trustworthy AI interactions in high-stakes environments.
Generative AI agents can reveal how personalization algorithms amplify toxic content in ways that vary dramatically by user ideology.
ANTAP achieves near-zero vulnerability to description-based attacks, fundamentally transforming how agents are evaluated and routed in multi-agent systems.
Grasp datasets can revolutionize robotic dexterity, enabling significant improvements in articulated tool use performance.
HTT enables tactile learning across diverse sensors, achieving adaptability that was previously unattainable in contact-rich manipulation tasks.
Estimating valid transport maps can be as hard as optimal transport, but under certain conditions, alternative maps can be learned with significantly higher accuracy.
Online imitation learning can outperform offline methods, but only when the student can effectively represent the expert鈥攔ealizability is key.
Despite holding privacy certifications, developers turn to Reddit for legal advice, revealing a critical gap in professional support for navigating privacy law.
Touch is not just an add-on; it fundamentally enhances object representation, leading to dramatic improvements in physical property estimation and manipulation tasks.
EFT enables LLMs to evolve solutions across diverse optimization tasks, achieving over 10% performance gains and state-of-the-art results in challenging mathematical problems.
A new AI tool can catch 34% more mathematical errors in scientific papers, transforming the peer review landscape.
Data mixing, especially with instruction-heavy data, emerges as the crucial factor for optimizing VLM training, challenging traditional filtering approaches.
EpiKV achieves 72% accuracy on MATH-500 with a 4096-token cache, rivaling the best attention-based methods while dramatically improving inference speed.
VibeAct reveals that leveraging real-time vibro-acoustic feedback can drastically enhance robotic dexterity in contact-rich environments, outperforming traditional proprioception methods.
The ABC framework empowers researchers with the largest open-source teleoperation dataset and a complete toolkit to accelerate advancements in behavior cloning for robotic manipulation.
ReStruct enables robots to adaptively steer their behavior in real-time, achieving unprecedented levels of task success and preference alignment without retraining.
Delayed start behavior can predict standardized test performance, revealing critical insights into student motivation and engagement.
A million p-bits in a single programmable architecture reveals a universal tradeoff between throughput and accuracy in distributed probabilistic computing.