From the 1 of 9 linked papers with an AI index.
9 papers
Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition
Bruce Coburn, Jingbo Yue, Jinge Ma +3
The paper presents Open-KNEAD, a training-free, locally run agentic system that breaks down meal images into individual food items, grounds each to a nutrition database, and improv…
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
Gautham Vinod, Bruce Coburn, Siddeshwar Raghavan +1
Accurate volume estimation of objects from visual data is a long-standing challenge in computer vision with significant applications in robotics, logistics, and smart health. Exist…
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
Gautham Vinod, Siddeshwar Raghavan, Bruce Coburn +1
Accurate dietary assessment is critical for precision nutrition, yet most image-based methods rely on a single pre-consumption image and provide only coarse, meal-level estimates.…
Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual Segmentation
Siddeshwar Raghavan, Gautham Vinod, Bruce Coburn +1
Audio-Visual Segmentation (AVS) aims to produce pixel-level masks of sound producing objects in videos, by jointly learning from audio and visual signals. However, real-world envir…
MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
Jinge Ma, Xiaoyan Zhang, Gautham Vinod +3
Food portion estimation is crucial for monitoring health and tracking dietary intake. Image-based dietary assessment, which involves analyzing eating occasion images using computer…
Implicit-Scale 3D Reconstruction for Multi-Food Volume Estimation from Monocular Images
Yuhao Chen, Gautham Vinod, Siddeshwar Raghavan +5
We present Implicit-Scale 3D Reconstruction from Monocular Multi-Food Images, a benchmark dataset designed to advance geometry-based food portion estimation in realistic dining sce…