Search papers, labs, and topics across Lattice.
15
3
13
13
The results suggest that gradient-based attribution can identify data responsible for subliminal learning in some settings, but that some approximations are more reliable than others.
ICON decomposition reveals the true reliance of deep models on concepts, debunking misleading correlations that traditional methods often overlook.
A relevance-based concept atlas reveals how tissue morphology directly influences spatial transcriptomics predictions, enhancing interpretability in pathology.
Sparsifying activations in collaborative inference may cut costs, but it exposes a hidden privacy risk from the positions of those activations that could enable re-identification.
Achieving formal privacy in federated learning without sacrificing model performance, FedKT-CSD outperforms traditional methods even under stringent privacy constraints.
Selective data sharing guided by XAI can significantly boost federated learning performance, achieving higher accuracy and faster convergence even in heterogeneous environments.
AIR achieves over 18% better perplexity than previous methods while retaining 60% of the parameters, revolutionizing LLM compression efficiency.
Outer informedness in INDEQS consistently reduces forecasting error on complex graphs, outperforming traditional methods even with similar parameter counts.
Activation probes can predict future behaviors in reasoning models with up to 91% accuracy, enabling effective steering without sacrificing output quality.
Steering isn't just a trick; it's a fundamentally different way to adapt language models, offering localized, reversible control that traditional fine-tuning can't match.
Activation steering turns interpretability into a hands-on debugging tool, but watch out for unintended consequences and limited generalization.
Forget IoU, measuring the structural compactness of attribution maps with Minimum Spanning Trees reveals fundamental differences in how models explain themselves.
Blockchain-based federated learning can be made practical by using multi-task peer prediction to overcome the computational bottleneck of contribution measurement.
Know where and by how much your PINN is wrong: a lightweight method provides pointwise error estimates without needing the true solution.
CLIP models exhibit surprising reliance on latent components encoding polysemous words, visual typography, and dataset artifacts, revealing hidden biases that can be amplified in downstream tasks.