food segmentation 1language injection 1large language models 1multimodal fusion 1transformer decoders 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion
Jui-Feng Chi, Wei-Ta Chu, Sheng-Long Lin
The paper proposes two plug‑and‑play modules that inject ingredient labels generated by large language models into image segmentation networks, boosting fine‑grained food image seg…
cs.CV2025
MVA 2025 Small Multi-Object Tracking for Spotting Birds Challenge: Dataset, Methods, and Results
Yuki Kondo, Norimichi Ukita, Riku Kanayama +21
Small Multi-Object Tracking (SMOT) is particularly challenging when targets occupy only a few dozen pixels, rendering detection and appearance-based association unreliable. Buildin…