Search papers, labs, and topics across Lattice.
DietDelta, a vision-language framework, is introduced to estimate food consumption at the item level using before-and-after images. The method uses natural language prompts to localize food items and estimate their weight directly from single RGB images, avoiding the need for depth sensing or explicit segmentation. Experiments on three public datasets show DietDelta outperforms existing methods, providing a new baseline for dietary image analysis.
Ditch the depth sensors: DietDelta uses before-and-after photos and natural language to pinpoint exactly what you ate, item by item.
Accurate dietary assessment is critical for precision nutrition, yet most image-based methods rely on a single pre-consumption image and provide only coarse, meal-level estimates. These approaches cannot determine what was actually consumed and often require restrictive inputs such as depth sensing, multi-view imagery, or explicit segmentation. In this paper, we propose a simple vision-language framework for food-item-level nutritional analysis using paired before-and-after eating images. Instead of relying on rigid segmentation masks, our method leverages natural language prompts to localize specific food items and estimate their weight directly from a single RGB image. We further estimate food consumption by predicting weight differences between paired images using a two-stage training strategy. We evaluate our method on three publicly available datasets and demonstrate consistent improvements over existing approaches, establishing a strong baseline for before-and-after dietary image analysis.