Search papers, labs, and topics across Lattice.
71 papers published across 7 labs.
ALE and SHAP emerge as the most reliable methods for understanding feature importance in heat demand forecasting, revealing critical insights for high-stakes AI applications.
Global sharpness regularization can lead to underfitting in Probabilistic Circuits, but a new adaptive approach recovers generalization while maintaining robustness.
DMDIntel reveals that LLMs can be interpreted with greater accuracy by leveraging dynamic mode decomposition, outperforming traditional methods in token attribution.
A hybrid concept bottleneck model boosts interpretability in cancer imaging diagnostics with only 10% annotation, achieving a mean concept AUC increase from 0.619 to 0.741 for mammographic masses.
The position where you measure a latent can distort your conclusions, with up to 11.9% of variance attributed to token selection rather than actual differences in the autoencoders.
Global sharpness regularization can lead to underfitting in Probabilistic Circuits, but a new adaptive approach recovers generalization while maintaining robustness.
DMDIntel reveals that LLMs can be interpreted with greater accuracy by leveraging dynamic mode decomposition, outperforming traditional methods in token attribution.
A hybrid concept bottleneck model boosts interpretability in cancer imaging diagnostics with only 10% annotation, achieving a mean concept AUC increase from 0.619 to 0.741 for mammographic masses.
The position where you measure a latent can distort your conclusions, with up to 11.9% of variance attributed to token selection rather than actual differences in the autoencoders.
The choice of prompt can skew model evaluation scores, revealing that two conflicting studies can be reconciled by simply altering the prompt used.
Engagement metrics of LLMs shift unexpectedly as failure rates decrease, revealing nuanced self-monitoring behaviors that challenge conventional assumptions about model reliability.
ALE and SHAP emerge as the most reliable methods for understanding feature importance in heat demand forecasting, revealing critical insights for high-stakes AI applications.
DECAF reveals that the trajectory of model responses can provide deeper insights than mere magnitude, with a staggering 96.4% alignment with observed behaviors compared to just 35% for traditional methods.
Uncovering stable behavioral patterns in blockchain transactions reveals both routine and malicious activities, transforming our approach to forensic analysis.
PRISM reveals that large language models exhibit cognitive specialization akin to that of aphasic patients, challenging our understanding of LLM interpretability.
SkillShapley reveals that not all steps in agent skills are created equal, enabling precise identification of high-impact actions that can enhance performance.
E2-Explainer reveals the hidden communication subgraphs that drive successful collaboration in LLM-based multi-agent systems, enabling both interpretability and cost efficiency.
Directly verbalizing sparse autoencoder features from LLM representations transforms how we interpret model behavior, making explanations more efficient and insightful.
Task progress in vision-language-action models can be read directly from their internal representations, even before task-specific training, revealing a surprising depth of interpretability.
Mr3D-VL outperforms existing models in cross-modal reasoning for brain tumor imaging, achieving a BERTScore of 0.856 and setting a new standard for interpretability in mpMRI applications.
Even state-of-the-art medical AI models like MedCLIP are vulnerable to real-world shortcuts that can compromise diagnostic reliability.
GDCE-I achieves faithful and interpretable counterfactual explanations for graph neural networks without compromising on the integrity of the data structure or the search space.
Conventional RoPE's non-repeating rotations may hinder its ability to access distant context, unlike periodic RoPE, which excels at recognizing modular languages.
Separating extraction from reasoning can achieve ten times fewer tokens while enhancing accuracy and robustness in AI responses to complex policy queries.
Localization of latent structures is achievable, but the subsequent gating and release mechanisms are fundamentally flawed, revealing critical limitations in model behavior adaptation.
JAPE reveals that modeling evolving dependency structures can significantly enhance anomaly prediction accuracy and explainability in multivariate time series.
Erasing task-vector interference in merged language models is directionally dependent, revealing that magnitude alone fails to capture the true dynamics at play.
Internal activation signals in LLMs can be manipulated to suppress implicit demographic influences more effectively than traditional prompting methods.
Saliency-guided cutouts can improve image classification for natural images but fail to enhance malware detection, revealing a critical domain dependency in training strategies.
The evolution of Class Activation Mapping reveals a shift from simplistic CNN explanations to sophisticated, multi-layered insights that challenge traditional evaluation metrics.
Executable roles derived from team trajectories can boost multi-agent language model performance by over 16 points, reshaping agent interactions.
HyperANFIS boosts predictive accuracy and rule quality by leveraging hyperbolic geometry, outperforming traditional ANFIS models.
The routing structure of a transformer can be induced by type-level supervision, yet it operates independently from the answer generation process, challenging traditional notions of model architecture.
Foundation models exploit subtle low-to-mid frequency cues to distinguish real images from diffusion-generated ones, challenging assumptions about semantic reliance in detection.
Groundedness Drift reveals that explanations can mislead in the presence of backdoor attacks, highlighting vulnerabilities in language model classifiers.
Accuracy and explanation faithfulness in malware detection are at odds, with the most interpretable models yielding only mid-tier performance.
Last-layer embeddings in protein language models often underperform, with shallow layers proving more effective for specific datasets like deep mutational scans.
The prediction direction in transformer models acts as a critical anchor, revealing a steep geometric gradient that shapes model behavior and task framing.
Mechanist uncovers a surprising safety risk where unsafe traits can transfer across modalities, challenging assumptions about training data safety.
Excess separability reveals that traditional methods for benchmark contamination detection can mislead, with real transformers showing significant variations in probe accuracy depth profiles.
Causal tracing reveals that crucial visual tokens in VLMs often come from unexpected areas, challenging assumptions about spatial localization in multimodal reasoning.
Visual signals, not language priors, drive attribute hallucination in VLMs, leading to a novel framework that effectively mitigates this issue.
Sparse autoencoders may misrepresent human category structures, failing to capture the nuanced typicality that dense embeddings can.
Truthfulness in NLP research has surged to 37% of papers by 2026, reflecting a critical shift in focus towards safety and alignment in generative systems.
Confident predictions in language models can be surprisingly fragile, and leveraging this fragility can dramatically improve error prediction and uncertainty estimation.
A phase transition measured to three decimal places reveals that common probing metrics may mislead researchers about a language model's capabilities.
Operationally useful explainable IoT intrusion detection hinges on a delicate balance of predictive quality, explanation cost, and stability, not just accuracy.
Coordinated multi-LLM reasoning boosts fault classification accuracy in automotive systems, achieving a remarkable 0.917 Top-1 accuracy while enhancing interpretability.
Achieving 90% accuracy in fault localization while providing interpretable explanations could revolutionize root cause analysis in automotive systems.
Steering persona features can amplify emergent misalignment rates in language models beyond what traditional fine-tuning achieves.
Achieving state-of-the-art predictive performance while providing transparent explanations, GARLIC transforms how we handle missing data in ICU time series.
UniProbe slashes object hallucinations in LVLMs by 55% during generation, all while operating with minimal latency.
Iterative erasure counts can mislead researchers about the true dimensionality of concepts in neural representations, revealing a deeper complexity in how we measure and interpret these dimensions.
MR-MoL reveals how GNN-derived attributions can transform molecular property predictions by making the reasoning process transparent and interpretable.
RoT explanations can provide critical insights into AI decision-making processes without requiring model access, making them invaluable for auditing and compliance.
High-FNL features boost model performance significantly, while surprisingly, low-FNL features are more effective for jailbreak mitigation.
BREAD outperforms traditional methods by delivering more accurate and faithful diagnoses of anomalies in AI systems, addressing a critical gap in anomaly detection frameworks.
Even when time-series models report the correct delay, they often ignore the relevant historical data, challenging assumptions about forecast reliability.
An entropy-centric approach to explainable AI significantly enhances the interpretability of semantic segmentation in remote sensing, outperforming existing methods.
Answers that maintain stability across different contexts are significantly more likely to be correct, revealing a new lens for assessing LLM reliability.
Forcing attention onto the output axis can degrade performance by up to 84 times compared to a random rotation, revealing a critical trade-off in transformer architecture.
Activation probes can uncover security vulnerabilities in AI-generated code that traditional prompting methods completely miss.
MMDiff reveals that multimodal SAEs can be powerful tools for both auditing and steering MLLM behavior, achieving up to 24% reduction in safety attack success rates without compromising performance on visual question answering.
Expert allocation in retinal pathology detection is not only disease-dependent but also enhances interpretability, revealing how models can effectively disentangle complex co-occurring conditions.
Simplifying financial decision-making models can enhance readability without sacrificing predictive power, challenging the notion that more complexity always yields better performance.
UNMASK reveals that automated discovery of spurious correlations can enhance model robustness without human intervention, achieving significant accuracy improvements on benchmark datasets.
Hierarchical probabilistic summaries can dramatically enhance out-of-distribution detection, outperforming traditional methods without needing extra in-distribution data.
Users can now interactively critique and optimize their AI-driven data queries, transforming opaque analytics into a transparent, controllable process.
A single adapter can unlock shared activation tools across different language models, enabling seamless integration and effective reuse without retraining.
Fake media consistently generates lower-magnitude representations, allowing simple statistical methods to rival complex detection systems.
Per-instance layer selection can significantly enhance language model performance, recovering nearly all benefits of an oracle while avoiding common pitfalls of global selection.
High confidence scores in diffusion language models can mask significant ranking failures, revealing a critical gap between internal error detection and external performance.
Emotion-sensitive neurons in LALMs are language-specific, but pooling cross-lingual evidence reveals powerful, transferable Multilingual Emotion Neurons that enhance affective control.
Explainability in IoT isn't just an add-on; it's a fundamental architectural requirement that can transform how we understand automated decision-making in connected environments.
Sparse autoencoders can be effectively leveraged for music retrieval, but traditional methods fail to capture the nuances of audio concepts—this new approach solves that problem.
SymDiag reveals that existing verification methods fail to diagnose reasoning errors effectively, providing a robust framework that localizes failures and generates actionable insights.