Search papers, labs, and topics across Lattice.
84 papers published across 4 labs.
Token-level feedback in SecOPD slashes adaptive prompt injection success rates from 94% to just 9%, revolutionizing defenses for AI agents.
ArmorOCR not only enhances adversarial OCR perception but also preserves competitive performance on standard OCR tasks, bridging a critical gap in model robustness.
Extracting hidden reasoning traces from black-box models is not only feasible but poses a substantial security risk, with EchoCoT achieving over 66% accuracy in retrieval.
TESTNAV achieves up to 2.15x faster exploration of compositional perturbation spaces while ensuring realistic and severe input failures are prioritized.
Subtle timing in subtitle presentation can amplify video jailbreak effectiveness against LVLMs, achieving unprecedented attack success rates.
Token-level feedback in SecOPD slashes adaptive prompt injection success rates from 94% to just 9%, revolutionizing defenses for AI agents.
ArmorOCR not only enhances adversarial OCR perception but also preserves competitive performance on standard OCR tasks, bridging a critical gap in model robustness.
Extracting hidden reasoning traces from black-box models is not only feasible but poses a substantial security risk, with EchoCoT achieving over 66% accuracy in retrieval.
TESTNAV achieves up to 2.15x faster exploration of compositional perturbation spaces while ensuring realistic and severe input failures are prioritized.
Subtle timing in subtitle presentation can amplify video jailbreak effectiveness against LVLMs, achieving unprecedented attack success rates.
Evolved ransomware can now elude detection for hours by optimizing encryption patterns through Genetic Algorithms, raising the stakes for cybersecurity.
Continual adaptation to evolving prompt injection attacks can enhance LLM defenses by up to 6.3 times, addressing a critical vulnerability in AI systems.
AEGIS slashes token recovery rates to nearly zero while preserving model performance, tackling all three critical channels of information leakage in federated learning.
Combining gradient obfuscation with image encryption can drastically reduce data leakage risks in federated learning without sacrificing model accuracy.
Chameleon slashes adversarial attack accuracy by up to 36.74% while minimizing bandwidth and time overhead, setting a new standard for Tor traffic defense.
Language models can reconstruct sensitive user data from innocuous outputs, with 2-digit secrets retrieved almost perfectly and 4-digit secrets at 82% accuracy, raising serious privacy concerns.
Post-editing attacks can destabilize watermark detection, but L-VQVAE's localized approach ensures reliable provenance in generative models.
Adapting four ImageNet-based models to new domains, the authors achieve up to 125x fewer evaluations while improving output quality by 29% on average.
Small, human-imperceptible image perturbations can drastically mislead Vision Language Models, revealing a significant security flaw in multimodal AI systems.
Backdoor detection just got a major upgrade—DistScan identifies attacks by analyzing shifts in prediction distributions, achieving a 27.32% accuracy boost over existing methods.
TestifAI can predict the robustness of deep learning models against complex perturbations with remarkable accuracy while slashing testing costs by up to 80%.
Attention transfer in Vision Transformers may look perfect, but it fails to enhance robustness under distribution shifts, revealing a hidden gap tied to training maturity.
EVADE uniquely enhances the reliability of medical VLMs by verifying diagnostic consistency across different image views, achieving up to 45% better calibration without retraining.
Failure-mode contextual bandits can boost model accuracy by over 4% on standard benchmarks while eliminating the need for additional human annotation.
Leveraging SMT conflict counts, SMTrap achieves unprecedented denial-of-service effects against large reasoning models without the need for model feedback or GPU resources.
Gradient Mirage disrupts the gradient-objective consistency that underpins gradient matching attacks, offering a robust defense without sacrificing model performance.
Covert coordination among language-model agents can be effectively monitored and mitigated without prior training on attack examples, achieving near-perfect detection rates.
A multi-agent architecture that prioritizes safety in connected vehicles can decisively mitigate the risks of false emergency alerts within a stringent 100-millisecond window.
Adversarial manipulations can selectively sabotage the learning process of CL networks, leading to catastrophic learning that undermines both new and retained knowledge.
FedLNS effectively screens out malicious updates in federated learning, achieving superior model performance without compromising client privacy.
Local LLM-based SSH honeypots can achieve superior shell emulation accuracy with the right prompting and fine-tuning strategies, but their effects can conflict in unexpected ways.
Compromising just one component of the BEC-Trap cipher collapses the entire keystream, offering unprecedented security in streaming encryption.
Solving the proposed one-way function requires tackling complex polynomial equations, making it a formidable challenge even for quantum adversaries.
Associative context retrieval can dramatically amplify the effectiveness of white-box attacks on LLMs without compromising their general performance.
OOD detection methods can significantly improve the reliability of EEG machine learning models in high-stakes environments, but their effectiveness is often underestimated.
Energy-efficient model cascades can falter under data perturbations, revealing hidden vulnerabilities that could undermine their effectiveness in real-world applications.
Targeted memory manipulation in LLM-agent communities can escalate group polarization, revealing a new vector for social influence attacks.
GraphGAN not only detects DDoS attacks more accurately but also generates realistic synthetic traffic to combat class imbalance, transforming the landscape of network security.
Learning to avoid pseudo-robust features can drastically enhance model performance on unseen adversarial examples.
Authenticated actors can exploit cloud notification systems to send unauthorized messages, bypassing traditional email security measures.
ADAPTD reduces false evictions while effectively containing attackers, proving that efficient threat defense can be achieved without sacrificing system performance.
NGS-Marker effectively thwarts partial infringement in 3D Gaussian Splatting, ensuring robust copyright protection for 3D assets.
Misleading attacks can exploit security copilots even with factually correct documents, revealing a critical vulnerability in RAG systems.
HarnessRisk reveals that up to 80.9% of adversarial attacks can succeed in agent harnesses, even when risk detection is high.
Reflex-Guard filters harmful prompts with 95.9% accuracy in just 37.6 ms, revolutionizing real-time safety for LLMs.
Communication costs for consensus protocols can skyrocket under certain adversarial structures, revealing surprising dependencies on termination requirements.
A content-agnostic leakage gauge can reliably detect context-leakage risks in LLMs, achieving near-perfect accuracy across diverse models and attack scenarios.
Realistic attack synthesis using tabular diffusion models can boost DDoS detection accuracy in 5G systems from catastrophic failures to perfect scores.
BULLSEYE achieves a staggering 9.5x to 72.5x reduction in Time-to-Exposure for firmware vulnerabilities, setting a new benchmark in directed fuzzing.
All tested GUI agents are alarmingly vulnerable to environmental injection attacks, with success rates reaching over 66%, revealing a pressing need for improved safety measures.
Deepfake speech detection may achieve sub-1% error rates in controlled settings, but real-world performance falters dramatically due to unforeseen challenges.
A groundbreaking testbed that seamlessly integrates hospital IT and OT systems reveals vulnerabilities and defenses in real-time, paving the way for enhanced cybersecurity in healthcare.
Unconditional attacks reveal that quantum key agreement protocols are fundamentally flawed when classical communication is involved, jeopardizing their security.
MLLMs can be tricked into unsafe behavior by seemingly harmless prompts and images, but COMIC effectively mitigates this risk by focusing on the operation-target relationship.
Attackers can bypass defenses by cleverly fragmenting harmful tasks, rendering current stateful defenses ineffective in real-world scenarios.
Attack rankings shift dramatically when evaluated under shared target-call budgets, revealing hidden efficiencies in traditional methods.
PACE achieves a flawless safety record in DeFi transactions, eliminating unsafe executions while leveraging LLMs for complex financial operations.
Transformer-based models exhibit stark differences in handling distribution shifts, revealing vulnerabilities that could jeopardize autonomous driving safety.
Reasoning-capable models can significantly reduce the impact of misinformation in RAG systems without the heavy computational costs of isolation.
State information in LLM-driven agents can be exploited, turning their task execution capabilities into a potential attack surface.
Attack strategies can now evolve and adapt in real-time, leading to a staggering 48.6-point improvement against state-of-the-art models.
Adversarial examples can exploit reconstruction-based detectors, leading to a significant drop in detection accuracy even under real-world conditions.
Despite high gesture recognition rates, leading VLMs struggle with safety reasoning, exposing critical gaps in their operational reliability.
AdROD achieves superior adversarial robustness with only 1.6% of the parameter footprint of traditional HyperNetworks, making it a game-changer for real-time object detection in autonomous vehicles.
Targeted structural probes reveal that some neural audio watermarks can be erased with a single attack, while others remain impervious, highlighting a critical divide in watermarking effectiveness.
DSPrompt reshapes retrieval semantics to thwart adversarial attacks without altering the retrieval pipeline, achieving robust defense with less than 1% additional parameters.
A single authorization rule can effectively prevent memory leakage across different audience groups in language agents, ensuring that sensitive information remains compartmentalized.
Stripped of safety, language models can be tricked into generating misleading yet confident responses, with up to 90% of outputs being decoys under attack.
Skill compositions can exploit safety gaps, with CompoSkill achieving up to 83.3% success in forming risky chains from individually certified skills.
Adversarial manipulation of sensor data can be detected with surprising accuracy by leveraging transaction timing and behavior, even in the presence of plausible telemetry.
SFMI achieves a remarkable 92.48% accuracy in reconstructing target identities, setting a new benchmark for model inversion attacks on face recognition systems.
Signaling storms in 5G networks can be triggered by exploiting the RACH procedure, crippling legitimate user access and revealing critical vulnerabilities in the initial access phase.
Forged-reasoning attacks can completely bypass traditional defenses, but Proof-of-Execution Memory ensures that only verified actions are accepted, achieving zero successful attacks across tested models.
Covert channels in LLM traffic can leak private information without any malicious intent, revealing a surprising vulnerability in seemingly benign interactions.
Skill, not lineage, is the key determinant of success in trusted-monitor ensembles, with implications for how we build and evaluate these systems.
Extreme fidelity loss reveals critical vulnerabilities in long-horizon tasks that standard accuracy metrics overlook.
Temporal inconsistencies in Digital Twins can serve as powerful indicators of cyber physical attacks, achieving up to 98% detection reliability without needing labeled attack data.
Ingestion-time defenses against coordinated poisoning are fundamentally flawed, allowing attackers to manipulate retrieval systems with minimal effort.
Security vulnerabilities in embodied agents are more complex than previously understood, with critical gaps in defenses against long-term attack propagation and multi-agent interactions.
PANDA can verify the robustness of neural networks with millions of parameters in minutes, all while keeping model details private.
Indirect prompt injection can compromise AI systems like DeepSeek Harness, with attack success rates reaching up to 25.5% under certain conditions.
A black-box defense that reduces harmful outputs in text-to-image models by 37.7% without needing any model retraining or internal access.
Malicious developers can exploit the interaction between benign-looking wrappers and crafted metadata to manipulate AI outputs without altering model weights, revealing a significant security vulnerability in AI deployments.
Ruling parties are more susceptible to poisoning attacks in generative search engines, highlighting critical vulnerabilities in information access.
Evaluation choices can drastically change the perceived effectiveness of quantum machine learning models for power-system attack detection, revealing a critical vulnerability in benchmarking practices.
Price manipulation attacks in DeFi can be mitigated by a novel framework that reduces attack-induced price deviation from 55.56% to below 5%.
Targeted bit-flip attacks can cripple Vision-Language-Action models, reducing their success rates to zero with just a few flips in key layers.
ARENA uncovers vulnerabilities in large audio-language models that traditional text-based red-teaming methods miss, achieving near-perfect safety metrics across multiple systems.
APC transforms multi-agent security by ensuring that delegated authority is tightly controlled, effectively blocking all data-stealing attempts in real-world scenarios.