Search papers, labs, and topics across Lattice.
This paper introduces MUL-T, a lightweight transformer framework designed to decode spatial cellular architecture in multiplexed tissue images by treating tissue organization as a masked contextual prediction task over discrete cell tokens. Unlike traditional methods that rely on handcrafted features, MUL-T learns contextualized embeddings without task-specific supervision, enabling it to capture complex cellular interactions efficiently. Evaluated on multiple clinically relevant tasks, MUL-T outperforms classical feature-based approaches and matches the performance of larger models while using significantly fewer parameters and lower training costs.
MUL-T achieves state-of-the-art performance in tissue image analysis with a fraction of the parameters, redefining efficiency in spatial cellular architecture modeling.
Understanding tissue organisation in multiplexed imaging requires modelling both cellular phenotypes and their spatial context. Existing approaches typically rely on handcrafted features, such as marker intensity statistics or cell-type proportions, which often fail to scale or generalise across cohorts with heterogeneous marker panels. We introduce MUL-T, a lightweight transformer framework that reframes tissue architecture as a masked contextual prediction task over discrete cell tokens. By learning contextualised [CLS] embeddings without task-specific supervision, the model captures higher-order cellular interactions while remaining computationally efficient. We evaluate MUL-T on several clinically relevant downstream tasks, including core-level tumour pattern classification, patient-level grading, PD-L1 positivity prediction, and cross-dataset treatment response prediction. Across tasks, MUL-T consistently outperforms classical feature-based baselines and achieves performance comparable to a foundation ViT model, despite substantially fewer parameters and lower training cost.