Search papers, labs, and topics across Lattice.
This study evaluates the capacity of four retinal foundation models to generate clinically relevant images from latent representations, focusing on the preservation of demographic and clinical information during the synthesis process. The results indicate that while the generated images retain phenotype information and outperform traditional latent diffusion methods on various prediction tasks, they struggle to align with classifiers trained on real images, revealing a significant synthetic-to-real representation gap. These findings underscore the potential of foundation models for controllable retinal image synthesis while emphasizing the necessity for improved alignment with real-world data distributions.
Generated retinal images can inherit critical clinical information, but they may not perform well against real-world classifiers, exposing a crucial representation gap.
Medical foundation models learn latent representations of clinically meaningful phenotypes, yet their ability to support controllable image generation remains largely unexplored. We evaluate four retinal foundation models within the representation tokenizer framework and examine whether demographic and clinical information encoded in latent representations from foundation models is preserved during synthetic image generation. We show that generated representations and images faithfully inherit phenotype information when evaluated within their originating foundation models, consistently outperforming conventional latent diffusion on multiple downstream prediction tasks. However, these gains largely disappear when evaluated using classifiers trained on real images, revealing a previously uncharacterised synthetic-to-real representation gap. These findings demonstrate that foundation-model latent spaces provide a powerful substrate for controllable retinal synthesis while highlighting the need to better align synthetic representations with real-image distributions.