Search papers, labs, and topics across Lattice.
64 papers published across 0 labs.
Unvoiced speech regions boost deepfake detection accuracy by nearly 50%, outperforming traditional full-audio approaches.
Navigating the fragmented landscape of explainability tools just got easier with Virgil, a system that empowers both experts and non-experts alike.
Uncovering a rich atlas of physical concepts in a neutrino model reveals that interpretability can dramatically enhance angular resolution in particle physics tasks.
Pruning can significantly undermine sparse autoencoder performance, but activation-aware methods offer a robust alternative that preserves interpretability in LLMs.
ICON decomposition reveals the true reliance of deep models on concepts, debunking misleading correlations that traditional methods often overlook.
Navigating the fragmented landscape of explainability tools just got easier with Virgil, a system that empowers both experts and non-experts alike.
Uncovering a rich atlas of physical concepts in a neutrino model reveals that interpretability can dramatically enhance angular resolution in particle physics tasks.
Pruning can significantly undermine sparse autoencoder performance, but activation-aware methods offer a robust alternative that preserves interpretability in LLMs.
ICON decomposition reveals the true reliance of deep models on concepts, debunking misleading correlations that traditional methods often overlook.
A quantum-inspired model reveals three distinct driving profiles and offers real-time insights into driver behavior, enhancing both interpretability and practical application in traffic systems.
Bridging the gap between attribution and counterfactual explanations, this framework achieves unprecedented stability and reliability in time series model interpretability.
VINCENT achieves a remarkable mean motif recall of 0.826, significantly outperforming existing models in identifying key molecular regions driving drug synergy.
Orthogonal projection reveals that protein language model embeddings capture critical biochemical features, with their removal leading to significant drops in predictive performance.
CBMs can significantly boost human-AI collaboration accuracy, but only under the right conditions, highlighting the delicate balance between interpretability and trust.
Spectral analysis reveals that the ability to detect machine-generated text hinges on text length and generation style, challenging conventional detection methods.
The J-lens reveals that language models operate with a surprisingly sparse causal structure, concentrating energy in specific pathways to predict future outputs.
Merging models without losing performance is possible by disentangling task-specific features in high-dimensional space, yielding significant improvements even in conflicting scenarios.
LMSM reduces LLM vulnerability to malicious prompts by over 90% while preserving throughput, revolutionizing how we enforce security in AI systems.
Subtle image perturbations can significantly influence AI-generated and human images' separation in CLIP space, revealing a profound gap between AI and human visual understanding.
EigenCL reveals that embedding NDRE trajectory dynamics into contrastive learning can significantly enhance the accuracy and interpretability of crop stress diagnostics.
Statistically independent factors can exhibit anisotropic behavior that traditional methods fail to capture, revealing deeper complexities in representation learning.
Symmetries in neural networks can be tracked through parameter adjustments, but fixed directions fail to maintain this relationship post-training.
$\texttt{findr}$ achieves the dual goals of interpretability and predictive accuracy in credit risk modeling, revealing when traditional explanations hold and when deeper insights are necessary.
Traditional interpretability fails to predict task-critical mechanisms before training, but this new framework bridges that gap, enabling more effective fine-tuning strategies.
Identifying distinct disease trajectories in T2DM reveals critical insights into patient progression and comorbidity risk that could transform clinical management.
Explicit frequency-domain evidence can dramatically enhance LLM-based time-series anomaly detection, revealing insights that time-domain methods miss.
qshap reveals how individual features drive model performance, offering a deeper insight into GBDT effectiveness than traditional attribution methods.
Mechanistic control over data generation reveals hidden model dynamics, leading to more diverse and effective datasets that enhance downstream performance.
Heterogeneous expert families can significantly boost interpretability and predictive performance in machine learning models, adapting to local data structures more effectively than traditional homogeneous approaches.
MLPs leverage specialized neurons to achieve superior data efficiency by creating localized representations, contradicting the notion of a single global feature space.
RACE reveals that a scalable statistical approach can vastly improve our understanding of neuron behavior in LLMs while slashing computational costs.
A unified framework that makes feature attribution accessible to both experts and non-experts, bridging the gap in model interpretability tools.
Undisclosed inference-time steering can systematically bias LLM outputs, challenging the assumption that model weights alone dictate behavior.
Richer trace representations can dramatically enhance failure attribution in multi-agent systems, achieving new benchmarks in accuracy.
Models can achieve near-perfect accuracy in resolving conflicting cues, yet their internal mechanisms can differ dramatically, challenging our understanding of model interpretability.
Sophisticated reasoning in AI models creates hidden geometric signatures that can be detected even when traditional linear methods fail.
Counterfactual explanations can empower individuals to contest algorithmic decisions, but only if they are tailored to effectively reveal underlying errors.
Local explanations of process monitor predictions can now be derived from a causal framework, revealing event influences with unprecedented clarity and stability.
Achieving 91.1% accuracy in retinal disease classification using interpretable, ring-based vascular features challenges the reliance on deep latent representations.
Mutual visual attention can be transformed into interpretable engagement metrics that empower non-technical users to grasp complex social dynamics.
EMFE achieves 94.6% accuracy in malaria cell classification while remaining computationally efficient and interpretable, challenging the dominance of deep learning in this space.
Multi-VLM fusion not only boosts face recognition accuracy but also delivers richer, more robust explanations that enhance transparency and auditability.
Prompt learning transforms concept profiles so dramatically that only 16% of initial top concepts persist post-optimization, revealing a strong link between these changes and accuracy gains.
DAPF models excel in dementia detection but struggle to provide trustworthy explanations, revealing a gap between performance and interpretability.
The Qwen3-Omni model can reconstruct complex narratives from garbled audio inputs, revealing a hidden layer of reasoning that challenges our understanding of audio language processing.
Unvoiced speech regions boost deepfake detection accuracy by nearly 50%, outperforming traditional full-audio approaches.
IMNO achieves superior accuracy and stability in predicting long-term dynamics of dissipative PDEs by leveraging their low-dimensional structure, outperforming traditional methods.
Local distillation reveals patient subgroups in cancer gene expression data that traditional models overlook, all while maintaining the accuracy of complex black-box models.
Influence functions can now pinpoint the most impactful training samples without requiring labels, revolutionizing data attribution in scientific missions.
LUX achieves unprecedented alignment between generated captions and localized pathological evidence, drastically reducing clinical hallucinations in endoscopic analysis.
Spectral alignment in deep neural networks is not a universal outcome of training but a complex interplay of transport dynamics and cancellation effects that varies by layer and scale.
Extraction indices, not semantic labels, dominate the influence of activation steering, reshaping our understanding of model behavior across varying contexts.
Counterfactual reachability reveals surprising disconnects with classifier accuracy, challenging conventional wisdom about model boundaries.
SAGE achieves 87.96% accuracy in predicting postpartum depression risk using just 16 carefully selected features, highlighting the importance of psychological and socioeconomic factors over demographics.
Identifying minimal sub-tournaments reveals why candidates lose, offering a structured approach to understanding tournament outcomes that could transform decision-making processes in competitive settings.
Achieving 99.86% accuracy with a 0.136% false-accept rate, EGAMA-RC redefines malware triage by prioritizing risk and novelty over mere classification performance.
Spatial encodings in vision-language models only influence answers at deeper layers, revealing a nuanced transport mechanism that challenges existing assumptions about their functionality.
Spatial reasoning in VLMs operates on coarse object localization rather than precise boundaries, revealing a surprising disconnect between knowing where objects are and how they relate.
Class-conditional typicality maps derived from SVD can significantly enhance OOD detection in Vision Transformers without requiring additional training.
OmicSync not only clusters spatial multi-omics data but also provides a reliability score and detailed explanations for each assignment, transforming how we interpret complex biological datasets.
Traditional attention-mask checks miss critical causality violations, while our new audit method pinpointed 100% of failures across multiple models.
Even top-performing models can produce misleading explanations, revealing a hidden crisis in the trustworthiness of Explainable AI.
Restricting LoRA updates to a compact mid-network region can recover over half of the performance gains, challenging the notion that full-layer updates are necessary for effective fine-tuning.
Control latents can be transformed into actionable physical model ensembles, leading to a 23% reduction in tracking error without changing the underlying policy.
Compact finite-state machines can predict LLM agent failures with up to 94% accuracy, revealing that deployment harnesses shape behavior more than the models themselves.
Sharing high-fidelity explanations in fraud detection can expose systems to serious privacy risks, but DP-FedSHAP offers a way to maintain transparency without compromising security.
Current multimodal large language models are highly susceptible to situational illusions, revealing critical vulnerabilities in their reasoning and perception capabilities.
AstroPT reveals that galaxy properties emerge in a predictable order during training, offering a new lens to understand concept emergence in LLMs.
EXPL-FR reveals that you can achieve interpretable face recognition without any text training, using a simple adapter to bridge vision and language spaces.