Search papers, labs, and topics across Lattice.
This paper introduces V-REX, a novel approach to training veterinary vision-language models (VLMs) that challenges the prevailing notion that fine-tuning larger foundation models is necessary for domain specialization. By innovating across the entire VLM pipeline鈥攅ncompassing text tokenization, pre-training, grounding, and inference鈥攖he authors achieve superior performance in generating diagnostic reports for veterinary X-rays with significantly fewer parameters and resources. The results demonstrate that V-REX not only outperforms existing open foundation models but also enhances data utilization and training efficiency, marking a significant advancement in veterinary radiology AI applications.
Rethinking the VLM training pipeline allows for the creation of specialized models that outperform larger counterparts while using far fewer resources.
While generalist VLMs are expensive to train, creating domain experts is widely assumed to require fine-tuning increasingly large foundation models. We show that, in veterinary radiology, this assumption is misguided. By rethinking the entire VLM pipeline - from text tokenisation and pre-training to grounding and inference - we demonstrate that careful engineering can yield models that outperform much larger foundation models from scratch, without relying on any other data. Our approach introduces new strategies for generative pre-training and grounding that improve training efficiency, increasing data utilisation and downstream performance. Using only a fraction of the parameters, data, and compute of contemporary generalist models, we develop the first VLM capable of generating diagnostic reports for veterinary radiographs, surpassing open foundation models on this task by significant margin.