Search papers, labs, and topics across Lattice.
The paper introduces ScaFE, a novel approach for classifying pathological scars by leveraging a large language model (LLM) to generate deterministic feature programs that assess scar attributes from clinical photographs. This method addresses the challenges of limited expert-labeled data and the need for local data governance by executing these programs in a restricted environment, ensuring patient data remains local and decisions are reproducible. ScaFE outperforms the strongest baseline, BiomedCLIP, achieving 81.0% site-macro balanced accuracy on a dataset of 600 photographs, while also demonstrating significant data efficiency by retaining 72.0% accuracy with only 10% of the development data.
LLM-generated feature programs enable a data-efficient approach to scar classification, outperforming traditional methods while keeping patient data secure and decisions auditable.
Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and substantial acquisition variation across hospitals. End-to-end image models remain data-dependent, whereas sending photographs to a hosted vision-language model (VLM) may conflict with local data-governance requirements and yields decisions that are difficult to reproduce and audit. We introduce ScaFE (Scar Feature Engineering), which transfers clinical knowledge from a large language model (LLM) into deterministic, executable feature programs instead of asking the model to diagnose images. A web-enabled LLM retrieves clinical evidence and synthesizes programs that measure visually assessable scar attributes. Candidate programs execute in a restricted local environment, and only aggregate validation statistics and feature-level SHAP summaries are returned for iterative repair and refinement; raw images and patient-level outputs remain local. A lightweight Random Forest then operates on the resulting structured representation. On 600 photographs from three hospitals under leave-one-site-out evaluation, ScaFE achieves 81.0% site-macro balanced accuracy, exceeding the strongest baseline, BiomedCLIP, by 10.0 percentage points. With only 10% of the development data, ScaFE retains 72.0% balanced accuracy and an 11.8-point lead. Iterative refinement also raises the executable-program rate from 66.7% to 95.0%, with verified evidence for 91.7% of the final features. These results show that LLM knowledge can support data-efficient, cross-site medical image classification through local and auditable feature programs rather than direct VLM decisions.