Search papers, labs, and topics across Lattice.
73 papers published across 2 labs.
Achieving a 2.37% boost in Union accuracy over the best baseline, Robust CurveMoE reveals a new paradigm for balancing adversarial defenses across multiple norm constraints.
LLM-enhanced GNNs may boost performance, but they also expose sensitive information, revealing a troubling vulnerability to privacy attacks.
Stealthy attacks can now be countered effectively in linear models, improving resilience against adversarial data manipulation.
Forget-Retain Alignment Gap reveals that the structure of weight updates, not just their distance, is key to preventing LLMs from relearning forgotten information.
Prompt framing is the key factor driving jailbreak vulnerabilities in MLLMs, revealing significant inconsistencies in multimodal safety alignment.
Achieving a 2.37% boost in Union accuracy over the best baseline, Robust CurveMoE reveals a new paradigm for balancing adversarial defenses across multiple norm constraints.
LLM-enhanced GNNs may boost performance, but they also expose sensitive information, revealing a troubling vulnerability to privacy attacks.
Stealthy attacks can now be countered effectively in linear models, improving resilience against adversarial data manipulation.
Forget-Retain Alignment Gap reveals that the structure of weight updates, not just their distance, is key to preventing LLMs from relearning forgotten information.
Prompt framing is the key factor driving jailbreak vulnerabilities in MLLMs, revealing significant inconsistencies in multimodal safety alignment.
Diverse refusal prefixes can significantly bolster the stability of refusal mechanisms, making them more resilient against targeted attacks.
Self-evolving coding agents can unwittingly amplify malicious skills, with self-poisoning rates soaring to 86.7% under tailored conditions.
Subgraph construction, not GNN reasoning, is the critical failure point in KGQA systems, accounting for over 99% of performance collapse under adversarial conditions.
Attack success rates plummet as this self-evolving defense framework learns from each jailbreak attempt, adapting in real-time without any parameter tuning.
LLMs can now proactively correct user input errors, significantly boosting their accuracy and reliability in real-world applications.
Flipping just a few bits can inflate LLM output by over 5900%, revealing a critical vulnerability in Mixture-of-Experts architectures.
Explicitly defining threat actors could transform how we assess risks associated with the release of open-weight AI models.
Verdict-only evaluations can misrepresent the effectiveness of automated code reviews, with PRGuard revealing a 1.38x improvement in identifying actual vulnerabilities compared to existing methods.
A lightweight defense can significantly reduce backdoor threats in RTL code generation without the heavy cost of full model retraining.
Adversarial attacks can mislead code search tools by altering identifiers, reducing retrieval accuracy by up to 77% without changing code functionality.
ReDiR slashes attack success rates to under 8% by embedding trajectory-level safety insights directly into the action generation process of LLM agents.
SkillShield slashes malware-generation severity from 3.37 to 0.58, showcasing a new frontier in prompt-space security for coding agents.
LMSM reduces LLM vulnerability to malicious prompts by over 90% while preserving throughput, revolutionizing how we enforce security in AI systems.
Backdoor detection in self-supervised encoders can be achieved with remarkable accuracy by leveraging generative priors, reducing the need for prior knowledge about the encoder or attack strategy.
Phantom Navigator achieves covert and precise UAV redirection, overcoming the limitations of traditional attack methods that are often costly and unreliable.
Trust-aware monitoring can restore decision-making consistency in multi-robot systems even when faced with sophisticated localization spoofing attacks.
Small gyroscope perturbations can lead to significant and controlled UAV displacement, revealing a new vulnerability in flight control systems.
Extracting LLM assets from edge AI chips is feasible through laser voltage imaging, exposing critical vulnerabilities in current deployment practices.
A stealthy backdoor attack exploits batch-dependent behaviors in Vision MoE, remaining dormant during audits but activating with alarming success during deployment.
Adversarial robustness evaluations in finance can vary by over 700 times depending on the evaluation protocol used.
Supervised UQ ensembles can drastically improve LLM hallucination detection, achieving superior performance with as few as 100 labeled instances.
A staggering 72.9% of medical chain-of-thought rationales fail to influence model answers, raising critical questions about their role in clinical reasoning.
Backdoor attacks in MLLMs can be effectively neutralized without prior knowledge of the trigger, achieving near-perfect defense in most scenarios.
Security-oriented prompts may reduce invalid outputs but paradoxically increase the prevalence of low-severity vulnerabilities in LLM-generated code.
Quantum algorithms are essential for identifying high-correlation cryptanalytic approximations, as classical methods face exponential query limitations.
Sparse backdoor attacks can misclassify over half of triggered anomalies while keeping clean accuracy losses below 2%, revealing a stealthy threat in federated learning systems.
A dual-layer architecture that completely eliminates revocation attacks and blocks all prompt-injection attempts, ensuring safer interactions for LLM agents in web environments.
A unified model that adapts to the interplay between perceptual fidelity and prompt alignment can achieve state-of-the-art performance while revealing interpretable insights into human judgment.
Even the strongest Android GUI agents are universally vulnerable to runtime anomalies, revealing critical flaws in their robustness.
Certifiable randomness can now be achieved unconditionally against low-query-depth quantum adversaries, eliminating reliance on unproven conjectures.
StepGuard slashes attack success rates by over 77% while maintaining nearly all utility, setting a new standard for safety in LLM-based agents.
Cloned outputs from LLMs can be detected with over 87% accuracy using behavioral signatures, revealing significant vulnerabilities in prompt security.
Transferable adversarial examples can exploit vulnerabilities in federated learning, but a new defense mechanism shows how to turn this threat into a strength.
GNN-based NIDS can be significantly hardened against structural adversarial attacks through a novel adversarial training approach that mimics real-world attack scenarios.
Attnlocate reveals that LLMs can effectively guard against injection attacks by pinpointing malicious behavior-guiding instructions in real-time, achieving impressive detection metrics across diverse models.
No trading system architecture is inherently safe; adversarial signals can compromise decisions across all roles, revealing a critical vulnerability in multi-agent setups.
RAGSentinel can filter out adversarial documents with high precision, ensuring that retrieval-augmented generation systems remain robust against sophisticated attacks.
NeuronGuard slashes attack success rates to near-zero while maintaining task performance, revolutionizing LLM safety alignment.
Watermarking can expose hidden vulnerabilities in audio deepfake detection systems, leading to significant performance drops that vary by dataset.
RAG systems are not just enhanced by external knowledge; they are also vulnerable to targeted attacks that can compromise accuracy, privacy, and fairness.
Misanthrope outperforms traditional feature detectors by preventing the identification of people, thus safeguarding privacy while maintaining high image matching accuracy.
Psychological jailbreaks reveal a new frontier in LLM vulnerabilities, achieving an 87.3% success rate through multi-turn persuasion tactics.
Recovering up to 80 bits of the AES128 key under KPA is now feasible in practical timeframes, challenging assumptions about AES's security.
Achieving 99.86% accuracy with a 0.136% false-accept rate, EGAMA-RC redefines malware triage by prioritizing risk and novelty over mere classification performance.
Adversarial prompts can exploit Gumbel-based verification, nearly doubling information leakage and undermining static defense thresholds.
LLMs are more likely to comply with unethical requests when benign framing tokens overshadow cue-tokens, revealing a hidden vulnerability in their alignment.
NICWhisper reveals that electromagnetic emissions from network interface cards can effectively identify network threats, achieving over 80% accuracy without analyzing packet data.
Choosing the right codec for face image compression can mean the difference between a 6.9% and a 98% false non-match rate, depending on your byte budget.
AutoSaddler achieves up to 10% performance improvement in LLM agents by automatically optimizing harnesses based on execution failure signals.
Semantic Overlays can reduce prompt injection attack success rates from 34.8% to just 6.6%, all while keeping the model's output readable and accurate.
Slopsquatting risks are significantly mitigated with a two-layer detection system that achieves 76% hallucination-free code generation, even in adversarial contexts.
Even top-performing models can produce misleading explanations, revealing a hidden crisis in the trustworthiness of Explainable AI.
Stitch achieves a remarkable 31% boost in pivot attack detection accuracy while keeping false positives to an astonishing 0.006%.
Only 4-16% of security rules in CLAUDE.md align with built-in controls, exposing a critical gap in rule enforceability.
Tighter bounds on instance encoding invertibility reveal that deterministic encoders can be just as secure as their randomized counterparts, transforming our understanding of data privacy techniques.
TrustShift attacks can manipulate agent trust dynamics, achieving a staggering 69.5% success rate before being mitigated by a novel defense framework.
ROBBIN achieves nearly 90% attack success while preserving over 83% accuracy, making it a game-changer for reliable backdoor attacks across varying DRAM devices.
Model robustness against training-time data contamination varies dramatically, challenging the assumption that clean data performance predicts real-world reliability.
A single interaction is all it takes for an adversary to manipulate LLM memory systems, steering outputs toward a pre-specified target with alarming precision.
Some protocols are vulnerable to traditional Dolev–Yao attacks but remain secure against rational attackers, revealing a critical gap in conventional security assessments.
TEE-X achieves GPU-level inference latency for large vision models while ensuring robust security in edge applications, transforming the landscape of model deployment.
As the number of sampled outputs increases, the risk of selecting unsafe outputs becomes asymptotically certain, even with seemingly effective safety proxies in place.
Task-asymmetric miscalibration in on-device language models leaves developers blind to when their outputs can be trusted, with confident-wrong outputs indistinguishable from confident-correct ones.
DFL-C not only ensures global model consistency in decentralized federated learning but also outperforms existing solutions in the face of Byzantine attacks.
RAG-based defenses can significantly reduce package hallucination rates in LLM-generated code, but only if matched to the specific threat model and utility requirements.
AdaptPrint reveals hidden LLM identities with up to 92.1% accuracy, transforming the landscape of security assessment for black-box AI services.
AEGIS can effectively defend against Indirect Prompt Injection attacks while maintaining low latency and high utility, a breakthrough for LLM safety.
Token-level feedback in SecOPD slashes adaptive prompt injection success rates from 94% to just 9%, revolutionizing defenses for AI agents.