5 papers
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
Ketong Chen, Yuhao Chen, Yang Xue
Despite the rapid progress of Vision-Language Models (VLMs), their capabilities are inadequately assessed by existing benchmarks, which are predominantly English-centric, feature s…
FLAD: Federated Learning for LLM-based Autonomous Driving in Vehicle-Edge-Cloud Networks
Tianao Xiang, Mingjian Zhi, Yuanguo Bi +2
Large Language Models (LLMs) have impressive data fusion and reasoning capabilities for autonomous driving (AD). However, training LLMs for AD faces significant challenges includin…
FoodTrack: Estimating Handheld Food Portions with Egocentric Video
Ervin Wang, Yuhao Chen
Accurately tracking food consumption is crucial for nutrition and health monitoring. Traditional approaches typically require specific camera angles, non-occluded images, or rely o…
6D Pose Estimation on Spoons and Hands
Kevin Tan, Fan Yang, Yuhao Chen
Accurate dietary monitoring is essential for promoting healthier eating habits. A key area of research is how people interact and consume food using utensils and hands. By tracking…
Dietary Intake Estimation via Continuous 3D Reconstruction of Food
Wallace Lee, YuHao Chen
Monitoring dietary habits is crucial for preventing health risks associated with overeating and undereating, including obesity, diabetes, and cardiovascular diseases. Traditional m…