Search papers, labs, and topics across Lattice.
UC Berkeley's AI research lab. Pioneering work in robotics, RL, NLP, and computer vision.
100
0
0
Models that ignore context may seem robust, but they can fail spectacularly when the context is actually trustworthy.
Retaining the right contextual information can boost long-context modeling performance by over 20% compared to traditional methods.
Language model agents struggle with oncall RCA, achieving only 25.3% accuracy on realistic tasks, revealing a critical readiness gap for production environments.
Leveraging weak supervision from clinical reports boosts CT scan segmentation accuracy by up to 22% with minimal labeled data.
The first three reduced minimal Weierstrass coefficients of an elliptic curve can be precisely determined from its Frobenius traces and conductor parity, revealing a deep link to its isogeny class.
FIDAC reveals how interpersonal distance can be accurately quantified from video, transforming facial detection data into actionable insights.
TPIPS reveals that traditional metrics fail to capture the nuanced aspects of visual similarity, leading to a performance gap that can be bridged with context-aware evaluation.
Anchoring visual SLAM to real-world metrics could revolutionize depth estimation, eliminating scale ambiguity and enhancing navigation reliability.
FilmGPT redefines video editing by turning chaotic footage into polished sequences without generating new frames, leveraging learned cinematic grammar instead.
Language models may silently skew their answers based on their own values, leading to potential misalignment with user intentions and preferences.
Automatic harness evolution may not be the silver bullet for LLM performance it was thought to be, often lagging behind simpler scaling methods.
A novel retraction-free optimization method achieves global convergence on the Stiefel manifold, drastically improving LoRA fine-tuning efficiency for large language models.
Prismata cuts attack success rates dramatically while ensuring web agents can still perform their intended tasks without developer input.
FourTune slashes memory overhead by 2.25x while matching the performance of full-precision fine-tuning in diffusion models.
State-of-the-art Vision-Language Models fall short in real-world robotic applications, revealing critical gaps in their reasoning capabilities.
Rethinking LLMs through the lens of world literature could revolutionize how AI interprets and engages with diverse cultural narratives.
STL-shaped rewards lead to tighter velocity tracking and more stable training for quadruped locomotion, outperforming traditional hand-crafted methods.
Online imitation learning can outperform offline methods, but only when the student can effectively represent the expert—realizability is key.
Claw-like agents are vulnerable to severe security breaches, with malicious plugins achieving a 100% success rate in attacks.
Muon achieves rapid convergence in matrix factorization by avoiding slow saddle dynamics and enabling high learning rates, aligning weights in just two steps.
Synthesizing 48,000 interaction trajectories without human input enables a humanoid robot to learn complex loco-manipulation tasks effectively.
Chai uncovers over 100 cryptographic vulnerabilities, including a critical flaw in an SSL library used by billions, by transforming how we approach vulnerability discovery.
Solving Blackwell approachability problems via Gradient Equilibrium oracles reveals a deep connection between these two seemingly distinct optimization frameworks.
Asynchronous OPD can boost training throughput significantly while managing the challenges of stale data, transforming the efficiency of large language model fine-tuning.
D2D transforms conversational product search by cutting conversation times by nearly 30% while boosting accuracy and user satisfaction.
Training data diversity is the secret sauce that boosts agentic model performance, with OpenThoughts-Agent achieving a notable accuracy leap over existing benchmarks.
Racing replicas instead of the clock allows Ambulance to achieve unprecedented throughput and low latency in BFT systems.
Transforming sparse rewards into dense feedback can accelerate RL training by significantly enhancing policy learning efficiency without compromising optimality.
Grounded verification in TEXEDO enables humanoid robots to execute complex motions that are both semantically aligned with text prompts and physically feasible.
Libretto transforms symbolic music generation into a structured, editable process, enabling LLMs to create and revise music with unprecedented precision and control.
The traditional complexity of leverage-score algorithms is misleading; the real challenge lies in identification, not accuracy, allowing for a dramatic reduction in query complexity.
Coding agents can now autonomously refine robotic manipulation policies to achieve a staggering 99% success rate on complex tasks, revolutionizing real-world robotics.
VIMPO achieves superior performance in reasoning tasks by enabling fine-grained credit assignment without the complexities of a critic, redefining the landscape of reinforcement learning for LLMs.
Retraining generative models with different seeds can shift FID scores dramatically more than merely resampling, revealing a hidden layer of randomness in model evaluation.
SC3-Eval achieves a remarkable 0.929 Pearson correlation in evaluating robot policies, revealing critical insights into their real-world performance.
Local ordinances, often overlooked in legal AI, are now accessible at scale with the launch of LOCUS, enabling deep analysis of everyday regulations.
Kappa deflation reveals that LLM-as-a-Judge models may be overstating their discriminative abilities by up to 41 percentage points.
Human videos can now be transformed into actionable manipulation data for robots, overcoming traditional barriers in hand-object interaction estimation.
Spatial attention in VLMs is nearly irrelevant to accuracy, with self-consistency emerging as the true indicator of reliability.
Extracting action signals from 32,041 hours of human video enables CAIP to outperform leading vision encoders in robotic manipulation tasks by over 30%.
Tactile-reactive policies can boost robotic manipulation success rates by over 30% through innovative data collection and a new Mixture-of-Transformers architecture.
Evolved playbooks can boost vulnerability detection rates by over 6x and outperform dedicated commercial products, reshaping the landscape of automated security auditing.
RHO achieves a 45.0% success rate in robotic tasks, 2.5x higher than the best multi-turn agent, showcasing a breakthrough in real-time control efficiency.
A mere 1% of poisoned samples can flip classifier labels, leading to catastrophic false positives and negatives in jailbreak detection systems.
Achieving up to 88x efficiency gains, Taylor-Calibrate transforms the way we initialize hybrid linear attention models, drastically reducing the training burden.
VisualClaw slashes API costs by 98% while boosting accuracy, transforming how VLMs can operate in real-time environments.
FTP-1 not only excels on familiar tactile sensors but also achieves unprecedented success on unseen setups, redefining the potential for cross-sensor generalization in robotic manipulation.
Current AI agents excel in structured tasks but falter at generating novel insights and tackling open-ended scientific challenges.
State-of-the-art surgical robotics policies can be disrupted by adversarial attacks, leading to a staggering 61% drop in task success rates.
Adversaries can achieve complete control over robotic policies in real-time by exploiting visual conditioning vulnerabilities, turning them into remotely piloted instruments.
QGF achieves superior performance in reinforcement learning by optimizing policies solely at test time, sidestepping the instability of traditional training methods.
LLMs reveal surprising strengths and weaknesses in analyzing security logs, with performance heavily influenced by model design choices.
Determination provenance reveals how to quantify and analyze the ambiguity in data systems, transforming our understanding of data resolution costs.
High-quality dense rewards can elevate robotic manipulation success rates from 50% to near perfection, transforming how robots learn from their environments.
StreamForce achieves real-time video generation with physical control, outperforming traditional models in both responsiveness and realism.
Active exploration can dramatically enhance adults' ability to reason about complex causal relationships, but even with this advantage, they still struggle compared to simpler tasks.
LadderMan enables humanoid robots to climb ladders and manipulate objects with unprecedented robustness and adaptability in real-world scenarios.
Site4Drug revolutionizes drug target selection by autonomously recommending binding modalities based on comprehensive evidence, minimizing the risk of biologically occluded sites.
Achieving competitive computational efficiency in Hartree-Fock theory while allowing for flexible orbital locality could revolutionize molecular simulations.
Reward models trained only on success are fundamentally misaligned with human values, leading to dangerous over-rewarding of poor robot behaviors.
ToggleCCI adapts to unpredictable traffic patterns, delivering significant cost savings by dynamically switching between VPN and CCI based on real-time cost trends.
Evolving coding problems can restore meaningful evaluation metrics for frontier models, revealing their true capabilities and enabling self-improvement.
Masking stale observations can boost search agent accuracy, but only under specific conditions—too much masking can backfire dramatically.
Tactile sim-to-real just got real: a physics-grounded representation unlocks zero-shot transfer for complex dexterous manipulation tasks, even without ground truth sensor calibration.
Introspection Adapters, a promising approach to LLM safety, can be completely defeated by exploiting architectural symmetries.
Transformers can provably internalize chain-of-thought reasoning, matching the sample efficiency of explicit CoT while eliminating its inference overhead.
By parameterizing electronic structure calculations with both nuclear position and momentum, this work unlocks more accurate simulations of coupled nuclear-electronic motion, including effects like chiral induced spin selectivity.
Entangled photons let you selectively excite specific biexciton states in quantum dots, opening new doors for quantum control.
Widely used approximations for modeling hot-exciton relaxation in semiconductor nanocrystals can fail, but mapping surface hopping (MASH) offers a more reliable alternative.
Stop hand-tuning your retrieval pipelines: BRANE slashes costs by up to 89% while matching accuracy by dynamically configuring pipelines per query.
Adaptive evaluation exposes a substantial vulnerability gap, revealing that existing defenses may underestimate the capabilities of distillation attacks.
Curiosity-driven agents can escape local loops in 3D environments by remembering where they've been and building a persistent map of the world.
Camera pose, largely ignored in video LLMs, unlocks significant gains in spatial reasoning and even improves general video QA when used as a lightweight supervisory signal.
LLMs get *worse* at forecasting high-stakes events like epidemics and financial crises as they get more capable, because they aggressively extrapolate growth and overestimate tail risk.
Stop writing brittle log parsers: Sieve uses LLMs to directly query raw security logs with natural language, outperforming hand-coded scripts on complex investigations.
Adversarial clothing with non-overlapping visible-thermal patterns can reliably evade RGB-T detectors, even transferring across different fusion architectures.
Forget scaling laws – the real bottleneck in associative memory isn't storage, it's retrieval: forcing a single "winner" costs you a logarithmic factor in capacity compared to allowing a ranked list.
Retrieval-augmented LLMs are surprisingly vulnerable to memory poisoning via synonym substitution, a loophole that gradient-based defenses can't close.
YouTube's recommendation algorithm pushes Kyrgyz children towards Russian-language content, even when they signal a preference for their native tongue, effectively amplifying colonial influence.
LLM-powered query reformulation, a hot topic in IR, often fails to translate gains from lexical to neural retrieval, and bigger models don't always help.
LLMs struggle with structured 2D tasks when inputs are serialized into 1D, revealing a surprising performance gap compared to vision-augmented models that directly process the 2D layout.
Multi-agent LLM systems are leaving performance on the table by treating structured agent interactions as generic traffic; Pythia shows how to unlock substantial gains by exploiting workflow semantics at the serving layer.
Forget hand-crafted examples: this system automatically generates worked examples tailored to student errors by mining common code patterns.
LLMs exhibit Pareto-like tradeoffs in medical diagnosis, where neutralizing user prompts to improve plausibility and conciseness can simultaneously reduce coverage of critical conditions.
Kernel launch overhead is a bigger bottleneck than you think: GPUOS achieves up to 15.3x speedup by fusing operations at runtime.
The dream of universal representations across modalities may be just that: scaling up datasets and relaxing constraints reveals that models trained on different modalities learn rich, but fundamentally different, representations of the world.
Claim verification in peer reviews just got a major upgrade with Peerispect, a tool that highlights evidence directly in manuscripts for rapid assessment.
Current LLM detection methods in peer review are fooled by hybrid human-AI workflows, mistaking AI-written text for AI-originated ideas.
LLMs may learn shared syntactic dependencies even with limited data, but they're still data-hungry toddlers compared to humans.
Generate diverse, physically plausible, and language-annotated whole-body motion data for humanoid robots at scale with this new interactive web-based pipeline.
AI audit standards can fail to ensure responsible AI practices due to vague requirements and undefined terms, even while appearing compliant.
Agentic data science pipelines often reach falsely optimistic conclusions, but two simple sanity checks can expose these unsupported claims by testing if the agent can reliably distinguish signal from noise.
Unlock zero-shot generalization in robot manipulation by generating diverse, affordance-aware training data with 3D generative models and Vision Foundation Models.
Verifier-free evolution can now match or exceed the performance of verifier-based methods, while slashing API costs by 3x and boosting throughput by 10x, thanks to a clever model orchestration strategy.
LLM-powered simulations of societal behavior risk encoding and amplifying existing biases unless strict ethical preconditions are enforced.
Cut LLM cold starts from minutes to seconds by pre-materializing CUDA graph execution contexts, sidestepping brittle kernel patching and heavyweight checkpointing.
MoEs can be pruned more effectively by considering cross-layer redundancy, leading to significant performance gains compared to uniform pruning strategies.
Core excitons in NaF decohere in under 8fs, and polarization-controlled attosecond spectroscopy reveals that bright excitons have s-like symmetry while dark excitons have p-like symmetry.
Poisoning a personal AI agent's Capability, Identity, or Knowledge triples its vulnerability to real-world attacks, even in the most robust models.
Get 3x the imitation learning performance from your robot with just a few extra cameras.