Search papers, labs, and topics across Lattice.
This paper introduces CM-MAE, a self-supervised framework designed for cross-scenario representation transfer in vision-wireless applications, addressing the challenge of representation failure due to variations in deployment conditions. By employing a soft contrastive alignment loss that leverages similarities in beam-power profiles, CM-MAE enhances the learning of representations from RGB frames and wireless measurements without relying on traditional calibration methods. The framework achieves a notable improvement in transfer performance, with a linear-probe average accuracy increase from 24.88% to 29.49%, and fine-tuning reaching up to 78.69% accuracy on unseen scenarios.
A novel soft contrastive alignment loss enables significant improvements in cross-modal representation transfer, achieving up to 78.69% accuracy in challenging unseen scenarios.
Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geometry change. This paper presents CM-MAE, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer. The evaluated real-data model uses only RGB frames and the measured 64-beam received-power vector available in DeepSense 6G; it does not use ray-traced paths, calibrated depth, or beam-index labels during pretraining. Its central pretraining term is a \emph{soft contrastive alignment loss}. Instead of making the synchronized image--wireless pair the only positive pair, this loss builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives. A masked joint decoder provides the complementary local objective by reconstructing hidden visual patches and wireless angular clusters under modality dropout. After pretraining, a differential-rate fine-tuning rule lets a new fusion head adapt quickly while the encoders move slowly. Under a sequence-disjoint DeepSense 6G protocol, adding the soft alignment loss improves a matched linear-probe transfer average from 24.88\% to 29.49\%. Mild fusion fine-tuning reaches 77.38\% Top-1 accuracy on unseen Scenarios 6--8, and optional transductive normalization adaptation reaches 78.69\%. Since the fusion setting uses the contemporaneous 64-beam power vector at inference, these results should be read as representation-transfer diagnostics, not as proactive beam-prediction or reduced-sweeping claims.