Search papers, labs, and topics across Lattice.
This study explores the effectiveness of simple transformations, specifically linear mappings, in translating representations across nine diverse text embedding models. By employing metrics such as CKA, downstream transfer, fidelity, and retrieval, the authors reveal that while some compatible pairs of models can successfully share semantic structures, others fail significantly, highlighting the nuanced dependencies on model architecture, training objectives, and data distributions. The findings challenge the prevailing notion of latent universality in text embeddings, emphasizing that simple mappings do not universally apply across heterogeneous models as previously suggested.
Simple transformations don't universally translate across text embedding models, revealing critical compatibility issues that challenge existing assumptions in the field.
We investigate whether simple transformations can translate representations across heterogeneous text embedding models. Understanding how independently trained models organize semantic information is an enabler for AI-to-AI latent communication without decoding into human-readable text. Focusing on lightweight translators such as linear mappings, we test the literature hypothesis of latent universality in a realistic text setting beyond simplified benchmarks. Across nine embedding models differing in architecture, pooling strategy, and training objective, we evaluate compatibility using CKA, downstream transfer, fidelity, and retrieval. Simple translators recover meaningful shared structure and support transfer for some compatible pairs, but fail sharply for others. Compatibility depends jointly on architecture, training objective, pooling, and data distribution. Overall, the results show that heterogeneous embedding spaces are not universally related by simple mappings as often suggested in some literature.