Search papers, labs, and topics across Lattice.
This paper addresses the challenge of generating multi-view fisheye images for autonomous vehicle perception by introducing MIVIFI, a novel framework that utilizes cross-domain learning to enhance data efficiency. The authors adapt the SyntheOcc architecture for fisheye data and leverage Equirectangular Projections to bridge the gap between limited fisheye datasets and more abundant standard multi-view images. Experimental results show that MIVIFI achieves high-fidelity image generation and allows for structural modifications of scene content, significantly improving the training of visual perception systems in autonomous vehicles.
Bridging fisheye and standard perspectives, MIVIFI enables robust multi-view image generation that overcomes data scarcity challenges in autonomous vehicle training.
Achieving 360{\deg} coverage is critical for the visual perception systems of autonomous vehicles. Fisheye cameras offer a cost-effective solution by enabling full surround coverage with as few as two sensors. However, existing multi-view fisheye datasets are limited, and synthesizing rare corner cases typically requires computationally expensive 3D simulations, hindering the training. While generative models have achieved significant success in standard perspective imagery, their application to wide-angle distortion remains unexplored. In this work, we formally introduce the novel problem of multi-view fisheye image generation conditioned on volumetric semantic representations and present two distinct methods. We first propose SyntheOcc-FE, which adapts the SyntheOcc architecture to fisheye data. While effective, this method is constrained by the scarcity of fisheye datasets, which limits its generalization. To overcome these limitations, we propose our second method, MIVIFI (multi-view fisheye), which leverages cross-domain learning with Equirectangular Projections. By bridging the gap between dataset domains using KITTI-360 fisheye images alongside nuScenes multi-view standard images, our approach enables high-fidelity manipulation of scene content. This framework enables the structural modification of semantic occupancy inputs to introduce or eliminate specific actors and facilitates the rendering of diverse meteorological conditions and illumination scenarios absent in the limited fisheye datasets. Quantitative and qualitative experiments demonstrate that our methods achieve robust photorealistic multi-view fisheye image generation and highlight the specific advantages of our cross-domain strategy for handling data scarcity.