Search papers, labs, and topics across Lattice.
Adversarial testing of AI systems, jailbreaking research, prompt injection defense, and robustness evaluation.
#16 of 24
3
CalibForge reveals that adversarial calibration can transform terminal task generation, leading to up to 30.04 percentage points improvement in model performance.
Adding correctly labeled examples can paradoxically increase learning difficulty by a logarithmic factor, challenging existing assumptions about sample exchangeability.
Removing just a few companion cells can drastically alter single-cell annotation outcomes, revealing a hidden vulnerability in existing tools.
Residualizing CAN message features against their normal baselines enables unprecedented detection performance, even against sophisticated attacks that mimic legitimate traffic.
A zero-length nonce can compromise the security of GCM and GMAC, enabling attackers to forge messages by recovering the hash key.
A novel algebraic approach to model extraction cuts clustering time by over 90%, enabling efficient attacks on hard-label max-pooling CNNs.
Adversarial prompts can hijack VLM-controlled robots with a success rate of up to 29%, exposing a critical vulnerability in their operational integrity.
Trajectory-poisoning can turn untrusted experiences into trusted skills, embedding malicious behaviors into self-evolving agents with alarming success rates.
AMS reveals that safety training modifications can significantly alter the activation landscape of language models, impacting their compliance with safety protocols.
Label-free reliability in vision-language models has a computable blind spot that can be systematically characterized and detected.
Misleading historical data can corrupt over 30% of tool-calling decisions, but a new method can restore accuracy by effectively transferring reliable policies from teacher to student models.
Smart-home agents struggle to differentiate between real commands and misleading ambient noise, with traditional detectors and MLLMs both failing in complementary ways.
ARIA can achieve a staggering 94.5% success rate in implanting covert backdoors in customized LLMs while ensuring high task performance.
A structured threat model reveals critical vulnerabilities in Quantum-as-a-Service platforms, exposing attack chains that could compromise quantum computing workflows.
UCD reveals that even state-of-the-art segmentation models like SAM3 can be severely compromised by a single, cleverly crafted adversarial perturbation.
Ambient temperature fluctuations can be harnessed to create robust adversarial attacks that consistently deceive multimodal perception systems.
MMLMs can be made 99% safer against harmful multimodal inputs without sacrificing utility, thanks to a novel calibration approach.
Robustness to adversarial perturbations can dramatically change the sample complexity landscape, shifting accuracy dependence from linear to polynomial rates.
The integration of LLMs into EDA workflows could significantly amplify hardware vulnerabilities, but innovative defenses like split manufacturing may offer a pathway to secure chiplet systems.
Gradient Immunity can significantly hinder malicious fine-tuning efforts, keeping attack success rates at pre-release levels while enhancing safety without user intervention.
TwinIR can degrade HD map accuracy by nearly 9% while remaining nearly invisible to the human eye, posing a significant threat to autonomous driving safety.
GUARD reveals that probing action-head dependence on multimodal evidence can significantly enhance failure detection in VLA policies, outperforming existing methods in unseen task scenarios.
Targeted safety sampling can slash attack success rates in fine-tuned LLMs from over 59% to under 14% with minimal additional data.
Social cues can lead LLM safety panels to a staggering 100% false-alarm rate, revealing a dangerous flaw in majority voting mechanisms.
LoginTrap exposes a staggering 86% success rate for phishing-style attacks on LLM-based web agents, revealing a gaping hole in authentication security.
A novel IDS framework achieves near-perfect accuracy while revealing critical vulnerabilities in replay buffers that can be exploited by adversarial attacks.
Trident exposes a staggering 522% drop in defensive performance of DRL systems against adaptive threats, highlighting their critical vulnerabilities.
ColorFD achieves superior black-box physical attacks on remote sensing object detectors, outperforming traditional methods and maintaining effectiveness in real-world scenarios.
Adversarial attacks can degrade the efficiency of Vision Transformers, but MOAT ensures that performance remains nearly intact, limiting GFLOPs loss to just 3.4%.
Non-imperative syntactic structures can undermine safety alignment in large language models, exposing them to sophisticated jailbreaks.
VLMs can flip predictions in nearly half of cases due to simple changes in presentation order, revealing hidden vulnerabilities in clinical reliability.
A novel poisoning attack that cleverly disguises misinformation as conflict-minimizing updates, achieving unprecedented success rates against RAG systems.
Malicious instruction detection can be significantly improved by adapting adversarial training to the context of the task, leading to better robustness against evolving attack strategies.
Strings alone can dramatically enhance secret detection, achieving over 80% semantic retention with only a third of the context.
Parameter-efficient adaptations in public models can leak actionable structural information, with family leakage rates surpassing random chance across multiple architectures.
Adversaries can extract sensitive operational intelligence from public safety communications even when content is encrypted, revealing a critical flaw in LMR security standards.
A single manipulated search result can dramatically amplify the effectiveness of attacks on LLM-based search agents, revealing critical vulnerabilities in their evidence-gathering processes.
Obfuscation defenses are failing, with DeepInvert exposing vulnerabilities that allow for high-fidelity token recovery from supposedly secure embeddings.
Season redefines adversarial attack strategies by seamlessly integrating structural and textural updates, achieving unprecedented transfer success across diverse model architectures.
Coding agents are alarmingly susceptible to malicious skill files, with exploitation rates exceeding 95% in some cases.
Adversarial techniques traditionally used for attacks are being repurposed by content owners to safeguard visual assets against unauthorized use and enhance accountability.
Adversarial attacks can exploit input-adaptive optimizations in Vision Transformers, undermining their efficiency without sacrificing accuracy.
PIMiner achieves up to 86.7% attack success rate against unseen LLMs with just 10 queries, revolutionizing prompt injection red-teaming efficiency.
Sensitivity and causality in language models are anti-correlated, revealing that early-layer interventions can inadvertently harm downstream performance.
Event timing can be weaponized to bypass ECG monitoring systems, achieving up to 66.7% suppression of critical heartbeat classifications without altering the data itself.
Utility misspecification can lead to significant performance drops in RL, but this new framework ensures robustness against such deviations, enhancing real-world applicability.
MAFIA reveals that memory-augmented LLMs can be compromised with a staggering 90.7% success rate, even under rigorous auditing conditions.
SRAP achieves a remarkable trade-off, enhancing image fidelity while maintaining robust identity disruption against face-swapping attacks.
A smaller batch size and larger learning rate can lead to flatter minima in SAM, revealing a critical trade-off in hyperparameter tuning that impacts generalization.
Tiny input changes can destabilize UAV tracking models, revealing a new attack surface that undermines their efficiency and accuracy.