Search papers, labs, and topics across Lattice.

Stanford's Institute for Human-Centered Artificial Intelligence. Focuses on AI research, policy, and societal impact.
100
4
0
Photoexcitation triggers a rapid reorganization of Au-ligand interfaces, reshaping electronic structures crucial for photocatalytic efficiency.
FBG whiskers can enable underwater robots to perceive hydrodynamic trails with remarkable accuracy, mimicking the sensory capabilities of harbor seals.
AnaDiffusion allows for precise and controllable editing of 3D brain MRIs, achieving the lowest FID scores while preserving anatomical integrity.
LLMs can autonomously generate prompts that rival expert-written ones, but still fall short in accurately interpreting scientific context and discovering literature.
Tying every view of a digital artifact to its version can boost knowledge-work agent performance by over 12 points on critical tasks.
Collective dynamics of AI agents reveal surprising patterns: while communication boosts accuracy on objective tasks, it can lead to political bias in group opinions.
Regular medical check-ups and mental health stress are critical predictors of chronic kidney disease, revealing new avenues for early intervention.
Simulation-based pre-training can drastically improve the dexterity of robotic hands, outperforming traditional training methods with just a fraction of real-world data.
Language models struggle to consistently encode the current year, with associative and declarative representations diverging in their update responses.
Current AI-generated videos that mislead viewers are also the most challenging for existing detection systems to identify, revealing a critical vulnerability in misinformation defenses.
CARD achieves a level of realism in credit card discussions that outstrips traditional simulation methods, revealing the power of structured conversational planning.
Speed Tuning achieves over 2.4x speed-up in robotic manipulation tasks without the need for extra data collection, revolutionizing policy execution efficiency.
Conversational phase structures can be reliably observed in clinical encounters, but patient states are only partially recoverable, challenging assumptions about transcript-based inference.
Alienation in AI research isn't just a personal experience; it's a systemic issue that can be resisted through collective action and critical self-reflection.
Test-time self-correction can boost LLM accuracy by over 30% on challenging reasoning tasks without the need for external reward models.
SmartMage redefines 3D scene understanding by dynamically selecting modalities, achieving state-of-the-art results while minimizing irrelevant computations.
Extending context in conversations can significantly amplify the risk of LLMs promoting delusional behaviors, challenging assumptions about model size and reasoning capabilities.
Skill-switching accuracy in LLMs drops significantly on complex tasks, but a new training approach boosts performance from 34.4% to 68.4% on challenging benchmarks.
SWD achieves high-fidelity circuit extraction with less than 1% of the data used by traditional methods, revolutionizing interpretability in pretrained transformers.
Achieving up to 2,500x speed improvements, string2string Studio revolutionizes string-to-string analysis by making complex algorithms accessible and interactive in the browser.
Long-term interactions with language models could lead to significant cognitive and emotional changes in users that short-term evaluations miss entirely.
Robots can now learn to manipulate dynamic environments effectively from just one static demonstration, drastically improving performance and efficiency.
Robots can now adapt to drastic physical changes in real-time, maintaining stable locomotion even with a locked leg or added weight.
VLA models can achieve a 66% success rate in contact-rich tasks by addressing precision and force failure modes, a leap from the previous 41% benchmark.
CryptoProver can independently verify cryptographic libraries in under 12 hours, ensuring the integrity of critical code without altering its execution.
System prompts in commercial AI products are often a mixed bag, with 40% harboring instructions that can undermine user interests, revealing a critical gap in accountability.
Current MLLMs falter under context shifts, with a notable inability to balance answering and refusal rates, as revealed by the new MMOOC benchmark.
Naive pretraining of Q-functions may hinder performance, while a simple ensemble approach can lead to over 1.26x improvement in fine-tuning effectiveness.
Achieving near-centralized collision avoidance performance with 28.5% fewer synchronization events could revolutionize autonomous spacecraft operations.
Zero-shot optimization of robot designs leads to a staggering 70% reduction in tracking error, showcasing the potential of motion-conditioned co-design.
Coding agents can now access repository context 8.7x faster, dramatically improving efficiency in navigating evolving codebases.
Proprietary MLLMs may achieve high diagnostic accuracy, but they still struggle with reliable clinical reasoning, revealing significant gaps in their practical utility.
FIDAC reveals how interpersonal distance can be accurately quantified from video, transforming facial detection data into actionable insights.
Renormalization can redefine how we understand and address the sim-to-real gap in robotics by leveraging effective parameters that capture omitted dynamics.
Achieving nearly linear sample complexity for distribution learning is possible by leveraging hierarchical comparability in query families.
EmoScope redefines emotional image editing by enabling users to discover unique, context-specific editing strategies rather than relying on static templates.
Grounding linguistic explanations in mathematically interpretable units can boost the trustworthiness of AI in medical applications by enhancing both visual localization and reasoning quality.
LLMs are overzealous tutors, intervening too soon and too often, which may undermine true learning and cognitive engagement.
Achieving a staggering 250,000x speedup in hologram synthesis without compromising image quality could redefine the future of 3D display technologies.
Aurora's forecasts may be skillful, but they reveal a troubling disconnect from the underlying chemistry, raising questions about their reliability for environmental policy.
Quantum autoencoders can match classical anomaly detection performance while fitting seamlessly into the resource-constrained environments of future collider experiments.
Solar Open 2 outperforms its predecessors and competitors with a groundbreaking 1M-token context window, redefining the capabilities of large language models in agentic tasks.
Revealing robot motion in video models can transform how we predict and control robotic actions, achieving high fidelity with minimal training data.
Imperfect LLM detectors can paradoxically drive users to increase their reliance on LLMs, ultimately degrading output quality.
BrainNext achieves top-tier performance in neuroimaging tasks by leveraging self-supervised learning on a massive dataset, highlighting the power of large-scale pretraining in medical applications.
Memory-centric architectures could revolutionize mass spectrometry search speeds, achieving over 100x acceleration and unprecedented energy efficiency gains.
Hierarchy-aware training combined with anatomy-guided learning boosts LUS video classification accuracy while enhancing model interpretability.
A unified framework that transforms how we build and debug asynchronous robot programs, ensuring reproducibility and systematic debugging across environments.
Aligning the final outcomes of training rather than just the training process leads to a significant boost in model performance, with Inf-Match achieving state-of-the-art results across multiple benchmarks.
Scaling visuomotor context to 8K timesteps enables robots to master complex tasks and adapt in real-time, outperforming previous models by a staggering margin.
OrthoPilot outperformed seasoned orthopaedic experts in diagnostic reasoning, achieving a 10.6% increase in management success for complex musculoskeletal cases.
Pix2Act transforms complex 3D manipulation into a simpler 2D prediction task, leading to significant performance gains and robustness against camera variations.
Training dynamics of Transformers can be reduced to a low-dimensional manifold, revealing how inductive reasoning emerges from data statistics and model initialization.
Video LLMs can significantly improve their QA performance by integrating spatio-temporal evidence, bridging the gap between accuracy and visual perception.
AMP achieves millimeter-level precision in 3D manipulation by transforming action learning into a pixel classification challenge, drastically improving inference speed and success rates.
Achieving submicrometer thickness in liquid sheets opens the door to unprecedented insights into ultrafast interfacial dynamics.
SBR reduces cognitive fatigue while achieving a 54.1% task success rate, outperforming traditional methods in real-time kinematic retargeting.
FourTune slashes memory overhead by 2.25x while matching the performance of full-precision fine-tuning in diffusion models.
Unlearning shortcuts doesn't guarantee their complete removal; ART reveals that some associations can still be functionally restored, challenging existing evaluation methods.
FILTR achieves up to 30x speed improvements in bioinformatics algorithms while simplifying the implementation of complex recurrence relations.
Strong learning can be achieved with significantly fewer calls to weak learners by exploiting the structure of list-decodable codes.
Linear attention fails to capture spectral variations in graphs, but Graph Convolutional Attention achieves superior denoising by directly utilizing the graph spectrum.
Charge conditioning in ML force fields can drastically enhance predictive accuracy while maintaining computational efficiency, achieving remarkable reductions in error metrics with minimal data.
LLM-driven program synthesis can automate EEG feature engineering while ensuring interpretability and high detection accuracy.
Full-sovereign scaffolding not only boosts user sovereignty scores but also curtails privacy violations and manipulative behaviors in personal agents.
Discounted occupancy-ratio realizability alone can enable robust offline policy evaluation, eliminating the need for stringent completeness assumptions.
VLMs struggle with raw medical data, achieving only a 48.6% success rate in standardization, revealing a critical gap in their clinical applicability.
The ABC framework reveals how thoughtful design can transform digital health interventions from theoretical solutions into sustainable real-world applications.
Scaling LLMs significantly boosts social simulation accuracy in well-represented domains, but fails to enhance calibration for human cognitive biases.
PaperPilot transforms scientific literature search by enabling users to iteratively refine their search strategies through an interactive workflow, achieving a remarkable reduction in execution errors.
SuperFlex achieves unprecedented reconstruction accuracy for 3D point clouds by enabling deformable superquadrics to represent complex geometries robustly.
Achieving state-of-the-art 4D reconstruction, this method transforms monocular videos into high-quality dynamic 3D representations, even in challenging conditions.
QuasiMoTTo achieves up to 47% fewer samples while maintaining accuracy, challenging the conventional wisdom that independent sampling is necessary for effective parallelization.
Diffusion models can outperform autoregressive counterparts in medical report drafting while offering a unique any-order infill capability that enhances usability for clinicians.
Stealth biases in language models can be reliably detected using a novel distillation technique that amplifies hidden signals, transforming bias detection into a practical tool.
Memory management emerges as a high-leverage skill that can double or quadruple the performance of LLMs in complex tasks without altering their core action behaviors.
Learning from failures can boost agent success rates by over 6% without extra training, reshaping how we approach agent improvement.
Traditional models can't handle belief contraction effectively, but a new mechanism reveals the complexities of belief dynamics in response to uncertain announcements.
Freeform Preference Learning allows robots to be trained on nuanced human preferences, leading to a dramatic 38% improvement in manipulation tasks.
EQMs reveal that explanation quality can be quantitatively assessed, offering a more reliable indicator of forecasting accuracy than traditional methods.
LLMs generate stark and homogeneous stereotypes that distort human interpretations, revealing a dangerous "stereotype hallucination" that undermines their predictive validity in novel contexts.
Traditional recommendation algorithms falter when faced with LLM agents, revealing a surprising shift from personalization to mere structural pattern matching.
Synthesizing 48,000 interaction trajectories without human input enables a humanoid robot to learn complex loco-manipulation tasks effectively.
Learned stopping can significantly boost reasoning model performance in complex tasks, but its effectiveness hinges on the problem's characteristics rather than being a one-size-fits-all solution.
Adapting pretrained policies with just a modest multisensory dataset can enhance robot manipulation performance across diverse tasks without sacrificing prior knowledge.
Touch is not just an add-on; it fundamentally enhances object representation, leading to dramatic improvements in physical property estimation and manipulation tasks.
Models that write intermediate states significantly outperform those that only report final answers, achieving up to 91% accuracy in predicting outcomes from edited states.
Liability insurance could be the key to unlocking scalable AI legal services by balancing risk and accountability in unprecedented ways.
MoRE achieves a staggering 44 percentage point increase in deployment success rates by seamlessly integrating behavior mode redirection into policy weights, eliminating the need for inference-time adjustments.
Terminal-use agents are still far from achieving reliable general-purpose performance, with top models only scoring 65.8% on a new benchmark that spans diverse real-world tasks.
SSA achieves superior long-context inference by leveraging gist tokens, outperforming traditional attention mechanisms without the need for complex architectural modifications.
Policies trained in SimFoundry's automated environments achieve up to 40% higher success rates in real-world tasks by leveraging affordance-preserving scene variations.
Data mixing, especially with instruction-heavy data, emerges as the crucial factor for optimizing VLM training, challenging traditional filtering approaches.
Finding stationary points in non-convex landscapes can be achieved with significantly fewer queries using a novel quantum approach.
Distinct model capabilities reveal that relational context significantly influences mental health assessments, with Claude-3-Haiku and GPT-4o leading in classification and trigger detection, respectively.
Achieving a 2-orders-of-magnitude energy efficiency improvement, the new AIS framework outperforms prior analog generative models by up to 4x.
Pretraining through play can revolutionize how robots learn dexterous assembly, achieving 60% success in tight insertions with minimal contact clearance.
Uninformative mode probabilities in trajectory forecasting can be transformed into robust predictions with simple post-hoc adjustments, enhancing model performance without retraining.
None of the 18 multimodal large language models audited are order-invariant, with flip rates revealing a staggering sensitivity to input ordering that challenges current evaluation practices.
Voice AI systems can recognize emotional cues but consistently ignore them in decision-making, leading to dangerous misinterpretations.