Search papers, labs, and topics across Lattice.
This paper addresses the challenge of fashion complementary image generation (CIG) by introducing a multimodal language-grounding setting that utilizes free-form instructions to create garments that match a seed item based on user intent. By enriching existing benchmarks with varying levels of linguistic specificity, the authors validate the effectiveness of their approach through human evaluation and various image quality metrics. The proposed StyleFlow model, which integrates seed images and natural language instructions within a unified multimodal transformer, demonstrates superior performance in generating stylistically coherent garments while maintaining lower architectural complexity and inference costs compared to traditional methods.
Free-form instructions enable StyleFlow to generate fashion items that are not only visually aligned with user intent but also stylistically coherent, outperforming rigid template-based approaches.
Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must interpret language in visual context. Existing CIG benchmarks rely on rigid template prompts (e.g., "a photo of a skirt"), failing to reflect natural user queries and obscuring model behavior across levels of linguistic specificity. We introduce fashion complementary image generation with free-form instructions, a multimodal language-grounding setting where a model generates a compatible garment from a seed image and a natural-language instruction. To this end, we enrich three CIG benchmarks with low-, medium-, and high-specificity instructions generated by a vision-language model and validated by human annotators. We instantiate the task with StyleFlow, a Rectified Flow Matching model that jointly conditions on the seed image and instruction within a single multimodal transformer. Across image quality metrics, catalog-alignment analysis, ablations, and human evaluation, StyleFlow consistently produces instruction-aligned and stylistically coherent garments while reducing architectural complexity and inference cost relative to auxiliary-module approaches.