Search papers, labs, and topics across Lattice.
To overcome the brittleness of traditional RF-only sensing in complex vehicular environments, the authors formulate BeamTransFuser, a multi-modal beam prediction framework that hierarchically fuses camera, LiDAR, radar, and GPS data. The system integrates a generative feature-reconstruction module to synthesize latent representations when individual sensor feeds drop or fail in practical deployment. Evaluated on a real-world multi-modal V2X dataset, the approach consistently outperforms existing beam prediction baselines and maintains high accuracy under incomplete sensing conditions.
Autonomous vehicles can maintain millimeter-wave beam alignment even when critical sensors drop out by leveraging generative cross-modal reconstruction across camera, LiDAR, and radar feeds.
Integrated sensing and communication (ISAC) provides a promising foundation for beam prediction in future vehicle-to-everything (V2X) networks. However, existing sensing-assisted beamforming methods still rely heavily on radio-frequency sensing, which may become unreliable in complex vehicular environments. Meanwhile, the growing availability of heterogeneous sensors, such as cameras and LiDAR, offers new opportunities to improve beam prediction through richer environmental perception. Motivated by this, this paper proposes a multi-modal beam prediction framework for V2X networks. Specifically, we develop BeamTransFuser, a hierarchical Transformer-based architecture that progressively fuses camera, LiDAR, radar, and GPS observations for robust beam prediction. In addition, to handle possible missing modalities in practical deployment, we introduce a generative module that reconstructs missing modality features from the available observations. Experimental results on a real-world multi-modal V2X dataset show that the proposed framework consistently outperforms representative baselines, while the generative module further improves robustness under incomplete sensing conditions.