Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Large Vision-Language Models Get Lost in Attention
Gongli Xi, Ye Tian, Mengyu Yang +5
Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer…
cs.AI2026
FoodCHA: Multi-Modal LLM Agent for Fine-Grained Food Analysis
Woojin Lee, Pranav Mekkoth, Ye Tian +2
The widespread adoption of camera-equipped mobile devices and wearables has enabled convenient capture of meal images, making food recognition a key component for real time dietary…