4 papers
OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
Qian Jiang, Zhecheng Shi, Jingpu Yang +2
The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the…
UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition
Shuai Zhang, Zhecheng Shi, Zhuxiao Li +4
Semantic segmentation of large-scale 3D point clouds is crucial for applications such as autonomous driving and urban digital twins. However, the sparse sampling pattern of LiDAR a…
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
Mingxuan Cui, Jingpu Yang, Fengxian Ji +7
Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and fine-grained glyph-level stru…
Edge-ANN: Storage-Efficient Edge-Based Remote Sensing Feature Retrieval
Xianwei Lv, Debin Tang, Zhecheng Shi +3
Meeting real-time constraints for high-performance Approximate Nearest Neighbor (ANN) search remains a critical challenge in remote sensing edge devices, which are essentially fusi…