Search papers, labs, and topics across Lattice.
This study investigates the significance of mapping non-maximal probabilities to Gaussian mixture model (GMM) components in the context of S-JEPA encoder representations. By comparing the performance of the REAL SOFT method against two control strategies鈥擣IXED-RANDPERM and UNIFORM-TAIL鈥攁cross multiple experiments, the authors find that REAL SOFT consistently outperforms the controls in recovering the original GMM tail and enhancing spectral dynamics. These findings indicate that not only the numerical values but also the specific mapping of probabilities to GMM components plays a crucial role in shaping the learned representations of the encoder.
The mapping of non-maximal probabilities to GMM components significantly influences the performance of S-JEPA encoders, revealing that structure matters as much as values in representation learning.
S-JEPA uses soft Gaussian mixture model (GMM) posteriors instead of hard cluster labels to preserve uncertainty. It remains unclear whether the probability values alone are sufficient, or whether it also matters which GMM components receive the non-maximal probabilities. We test this with two matched controls. FIXED-RANDPERM keeps the top-1 component and probability together with the multiset of non-maximal probability values, but reassigns those non-maximal values using a mapping fixed for each physical frame. UNIFORM-TAIL keeps the top-1 component, its probability, and total non-maximal mass but distributes that mass uniformly. Across three independent seeds, REAL SOFT outperforms both controls on two frozen Encoder readouts. It provides better recovery of the original GMM tail and greater accessibility of spectral dynamics over short time scales after controlling for the complete spectrum of the current frame. In two exposure experiments, both readouts improved overall as more frames retained the original mapping. We also descriptively follow one Phase 2 trajectory after the switch to the online GMM. These results show that the numerical probability structure of the soft target does not fully determine the learned Encoder representation. The mapping of non-maximal probabilities to GMM components also matters.