Search papers, labs, and topics across Lattice.
This paper introduces TransPhy, a novel framework for visually grounded image editing that enhances the capability of existing visual in-context learning (VICL) methods by focusing on physically grounded transformations. By decomposing the process into physical-rule induction and transition-aligned rendering, TransPhy effectively predicts transformation rules and adapts them to specific scene contexts, leading to improved adherence to physical rules and better generalization to unseen scenarios. The results demonstrate a significant advancement over prior VICL approaches, with a benchmark that includes 74 transformation rules and over 5,200 image pairs for robust evaluation.
TransPhy achieves unprecedented accuracy in physically grounded image editing, outperforming traditional methods by improving rule adherence and generalization to novel transformations.
Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a source--target exemplar pair and a query image, physically grounded VICL requires a model to infer the demonstrated transformation, adapt its effects to the query-specific scene context, and preserve rule-irrelevant content. We introduce PhysVICL-74, comprising 74 physically grounded transformation rules and 5,240 source--target image pairs that form nearly 75K training and evaluation contexts. Its benchmark split separately evaluates novel-instance transfer and unseen-rule generalization. We further propose TransPhy, a framework that decomposes physically grounded VICL into physical-rule induction and transition-aligned rendering. TransPhy first predicts the demonstrated rule and an explicit query-specific target-state description, and then synthesizes the target image through token-wise mixture-of-experts adaptation, with expert routing guided by localized transition cues. Experiments show that TransPhy improves physical-rule adherence, query consistency, and unseen-rule generalization over existing visual in-context editing methods.