Search papers, labs, and topics across Lattice.
82 papers published across 7 labs.
Some World Action Models can be steered to enhance robustness against perturbations without additional training, revealing a new avenue for improving model reliability.
Interleaved noise injection not only boosts model robustness but also enhances performance on clean and out-of-distribution data without added computational cost.
Causal diagnosis transforms robot action testing from a blind resampling process into a targeted, efficient strategy that reduces failures by over a third.
Deception in language models can be traced to specific internal features, revealing a pathway for proactive security measures against malicious behaviors.
Existing memory defenses fail to protect against complex adversarial attacks, exposing LLMs to persistent vulnerabilities that can distort behavior over time.
Some World Action Models can be steered to enhance robustness against perturbations without additional training, revealing a new avenue for improving model reliability.
Interleaved noise injection not only boosts model robustness but also enhances performance on clean and out-of-distribution data without added computational cost.
Causal diagnosis transforms robot action testing from a blind resampling process into a targeted, efficient strategy that reduces failures by over a third.
Deception in language models can be traced to specific internal features, revealing a pathway for proactive security measures against malicious behaviors.
Existing memory defenses fail to protect against complex adversarial attacks, exposing LLMs to persistent vulnerabilities that can distort behavior over time.
Automating the synthesis of leakage contracts could revolutionize CPU security by eliminating the need for extensive manual effort in their development.
Evidence-grounded detection in FlowGuard reveals that traditional semantic analysis can miss critical execution-related risks, achieving up to 2.23x faster scanning.
Suppressing certain clinical covariates can actually improve the predictive power of prostate MRI grading models, revealing hidden dependencies that could mislead performance assessments.
Adversarial examples in vision-language models can be detected by their tendency to stray further from the data manifold, revealing a critical vulnerability in multimodal AI systems.
Random Logit Scaling can drastically reduce adversarial attack success rates while preserving model accuracy, challenging the effectiveness of current black-box defenses.
Poisoning pretraining data through public discussion interfaces poses a significant threat, with the potential for undetectable harmful behaviors in language models.
LLMs can misinterpret benign text as physically dangerous actions, but a new probing method achieves over 99% accuracy in identifying these risks without relying on explicit unsafe keywords.
DataShield reveals that aligning consensus subspaces across multiple LLMs can drastically enhance safety by filtering out risky fine-tuning data more effectively than previous methods.
Traditional sandbox detection methods falter against AI-capable malware, exposing a critical vulnerability in current security measures.
GlobalForge shifts the focus from fragile local artifacts to robust global structures, achieving a 5.89% accuracy boost in detecting AI-generated images under real-world conditions.
Targeted illumination attacks can reduce VLA model task success rates to zero, exposing a critical flaw in current defense strategies that misinterpret color information.
Malicious nodes can exploit blocklace's design to overwhelm correct nodes with an infinite stream of arbitrary updates, threatening system integrity.
Automated adversary emulation can achieve over 84% execution success by intelligently revising playbooks based on real-time failures.
Introspective attention modulation can significantly enhance the safety of T2I models without sacrificing quality, outperforming existing methods like concept erasure.
ARMOR++ reveals that leveraging semantic priors can dramatically enhance the effectiveness of adversarial attacks on deepfake detectors, achieving unprecedented transferability across diverse models.
Adversaries can achieve up to 88.2% success in exploiting LLMs for log analysis through passive prompt injection, highlighting a critical security vulnerability.
Simplifying the attack pipeline for VLPMs can lead to a remarkable boost in transferability, outperforming complex methods with less resource consumption.
Attackers can weaponize ordinary project documentation to compromise AI coding agents, exposing a critical security gap in how these systems handle package installations.
Existing payloads in memory can compromise future agent behavior, revealing a critical vulnerability in memory-based AI systems.
Offensive security agents can significantly outperform proprietary systems when evaluated through a cost-aware lens, while defensive agents reveal a stark reliance on tool discipline over sheer computational power.
Safety in imitation learning can be achieved even in the face of significant distribution shifts, as shown by a UAV navigating uncertain environments while avoiding danger zones.
Adversaries can now exploit AI systems by manipulating behavior rather than compromising infrastructure, necessitating a radical rethink of penetration testing strategies.
Search agents can fail dramatically when faced with unreliable evidence, revealing substantial performance disparities that traditional benchmarks overlook.
Vulnerabilities in reusable agent skills can emerge at every stage of their lifecycle, not just during execution, revealing a critical oversight in current security practices.
Transaction privacy in encrypted mempools can create exploitable economic gaps that amplify vulnerabilities in perpetual funding mechanisms.
Elton can prove error bounds on security properties that were previously unaddressable, revolutionizing how we reason about adversarial probabilistic programs.
Latency differences in cloud configurations can collapse anonymity from four-way to three-way against adversaries, revealing critical vulnerabilities in Moving Target Defense strategies.
Exact verification of ReLU networks remains intractable even with random parameter noise, challenging assumptions about the feasibility of polynomial-time verification methods.
Generating text outside the detector's training distribution can defeat even the most advanced adversarial fine-tuning techniques, revealing a critical vulnerability in current detection systems.
Whistleblower protections can achieve robust privacy guarantees that significantly outpace traditional methods, ensuring safer reporting environments.
User-level permissions in AI agents are not just a feature; they are essential for mitigating risks like unauthorized transactions and data leaks.
Encoded adversarial prompts can exploit AI safety mechanisms with alarming effectiveness, revealing significant vulnerabilities in current models.
Natural alignment faking is prevalent in advanced models, with significant implications for the reliability of compliance under monitoring.
Traffic-Aware Randomized Smoothing boosts certified accuracy of LLM-based IDS by up to 72 percentage points compared to traditional methods, revealing the critical importance of aligning noise injection with attacker-controllable features.
Near-term quantum machine learning can expose weaknesses in post-quantum cryptography, potentially undermining its security assurances.
Frontier AI agents can autonomously conduct clinical security audits, but not all models perform equally, revealing significant efficiency gaps.
By constraining reward functions to Control Barrier Functions, this approach achieves safe adversarial imitation learning that adapts directly from expert observations without requiring labeled data.
Every randomized Byzantine Agreement protocol faces a fundamental round complexity barrier that scales quadratically with the number of corrupt parties.
Isolation is the key to understanding and mitigating the systemic failures of LLM-agent interactions, revealing how vulnerabilities propagate across multiple boundaries.
Sequence models can significantly outperform traditional approaches in PV power forecasting, especially under high levels of weather prediction uncertainty.
Temporal editing can exploit vulnerabilities in handwriting recognition systems more effectively than traditional image-based adversarial attacks, achieving superior transferability with minimal visual distortion.
Adversarial robustness, not rote memorization, is the hidden culprit behind training data exposure in image reconstruction attacks.
Adversarially-trained models may sacrifice up to 29.5 percentage points in clean accuracy compared to their vanilla counterparts, challenging the notion that robustness comes without significant cost.
Static deepfake detectors are failing in the wild, but a continuously evolving system like BMF can achieve AUC scores that surpass even the best commercial solutions.
A novel trust model for language models dramatically boosts resistance to untrusted input manipulation while maintaining high-quality output.
LLMs can inherently recognize policy violations, and PVDetector exploits this to achieve unprecedented detection accuracy against prompt injection attacks.
JADR reveals that quantization can dramatically alter a model's internal safety mechanisms, challenging assumptions about robustness in LLMs.
Adversarial binaries can manipulate LLMs into making unsafe proposals, but robust authorization controls can effectively prevent this exploitation.
Every configuration of LLMs and agents hallucinated skill names, with rates averaging 36%, revealing a critical vulnerability that could be exploited for supply-chain attacks.
TrustVLA can detect and recover from VLA backdoors with minimal clean data, all while preserving normal task performance.
Revealing that the Gale-Shapley algorithm can expose all honest participants' preferences under specific conditions raises critical concerns for privacy in matching applications.
Adversarial attacks on quantum classifiers become exponentially more costly as model complexity increases, revealing a built-in defense mechanism against gradient-based strategies.
A detection framework for GNSS spoofing achieves over 95% accuracy in identifying timing anomalies, crucial for maintaining TDD network integrity.
Just one terminal bit per string can identify any countable collection of infinite languages, radically simplifying language learning models.
HyperSafe slashes harmful response rates in fine-tuned language models to below 1% without sacrificing task performance, revolutionizing safety alignment strategies.
Once a capability is discoverable, autonomous exploitation becomes trivial, challenging our understanding of security vulnerabilities.
Adversarial attacks on vision-language agents reveal critical vulnerabilities, with multi-view optimization strategies proving significantly more effective than isolated approaches.
Turn-level credit assignment can boost jailbreak success rates to over 98%, revealing the limitations of traditional trajectory-based methods.
Front-running can be effectively neutralized by a power-weighted randomized lottery that reshapes transaction incentives and preserves user rewards.
Evolved programs can exploit perceptual hash algorithms with unprecedented efficiency, exposing significant vulnerabilities in content moderation.
Existing security scanners misidentify nearly half of MCP server risks, challenging the reliability of current security assessments in LLM applications.
A staggering 68.8% of LLM-generated code snippets contain security vulnerabilities, revealing a complex web of interrelated flaws that challenge traditional isolation strategies in software security.
Local monitors can be fooled by distributed backdoors that split harmful payloads, exposing a fundamental flaw in multi-agent safety systems.
Input-aware dynamic backdoors can exploit QNNs with unprecedented stealth and specificity, posing a serious threat to their secure deployment.
AHA reveals a reusable vulnerability core across production LLM agents, significantly enhancing the efficiency of red-teaming efforts.
AMT-X reveals that existing safety evaluations can miss up to 33% of actionable harm by conflating partial success with full operational capability.
LLM translations of adversary emulation plans can retain structural integrity but often miss critical telemetry details, leading to operational gaps.
Gauntlet outperformed human reviewers in technical critique of computer architecture papers, revealing that LLMs can achieve significant analytical depth through structured multi-agent collaboration.
Outlier events can be a powerful tool for falsifying causal graphs, revealing inconsistencies that traditional methods might miss.
Gate fingerprints reveal that decoder systems in quantum computers can unintentionally leak critical information about ongoing computations, posing a significant security risk.
Aggressive path randomization can cut attack success rates to as low as 4-20% while boosting network throughput by nearly 31%.
Naive multimodal fusion fails dramatically under polluted conditions, revealing that selective trust can significantly enhance QA performance.
SAC-driven attacks can stealthily escalate maintenance costs in industrial systems by manipulating control signals without triggering detection.
Watermarking can effectively safeguard LLMs from model stealing attacks without sacrificing performance.
Indirect data poisoning can enable scientific fraud at an unprecedented scale, with a staggering 49.56% success rate in corrupting AI-driven research outputs.
PromptGraph reveals that modeling contextual relationships between text spans can dramatically enhance privacy preservation in LLM inference without sacrificing utility.
AWM enables autonomous vehicles to effectively learn from adversarial scenarios, significantly enhancing their robustness in rare and critical traffic conditions.