Search papers, labs, and topics across Lattice.
This study investigates the application of Dowker homology, a topological tool, to analyze sentence similarity by treating token embeddings from transformer models as point clouds. The researchers demonstrate that Dowker homology effectively captures sentence similarity information, validated through regression against ground-truth similarity scores, and provides a means for visual inspection of similarity data. While they derive single-number summaries from Dowker homology for practical use, these summaries do not outperform traditional sentence similarity measures based on established pooling methods.
Dowker homology reveals nuanced insights into sentence similarity, but traditional methods still hold the edge in performance.
Dowker homology is a topological tool that may be used to analyze the relative position of two point clouds living in a common space. We investigate whether Dowker homology captures sentence similarity information by treating the embeddings of the tokens that constitute a sentence pair as a pair of point clouds in the latent space of a transformer model, using both models that have and have not been fine-tuned for sentence similarity. We find that Dowker homology captures sentence similarity information, as measured by regressing Dowker homology features onto ground-truth similarity scores, and that it can be used for visual inspection of similarity data and models. In an attempt to make Dowker homology readily applicable, we derive from it single-number summaries that we expect to capture sentence similarity directly. These turn out to work reasonably well, but without outperforming standard sentence similarity measures based on established pooling methods.