Search papers, labs, and topics across Lattice.
RelightFormer adapts a video foundation Transformer into a feed-forward generative model for direct single- and multi-view object relighting, entirely bypassing traditional inverse rendering and explicit intrinsic decomposition. By leveraging a latent illumination module with cross-attention and permutation-invariant positional encodings, the network conditions spatial features on target environment maps while enforcing view consistency across unordered camera perspectives. Trained on the newly introduced Laval Objaverse Dataset containing 90K objects and 39K lighting environments, it delivers state-of-the-art photorealistic relighting and zero-shot generalization across single-, multi-, and novel-view benchmarks.
Explicit intrinsic decomposition is no longer a prerequisite for multi-view relighting: video foundation architectures can directly infer complex light transport and 3D geometry in a single feed-forward pass.
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.