7 papers
OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
Qian Jiang, Zhecheng Shi, Jingpu Yang +2
The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the…
LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents
Jingpu Yang, Fengxian Ji, Zhengzhao Lai +8
Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic laboratory environments remains challeng…
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
Mingxuan Cui, Jingpu Yang, Fengxian Ji +6
Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and fine-grained glyph-level stru…
FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
Fengxian Ji, Jingpu Yang, Zirui Song +5
Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluations offer limited coverage,…
FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
Fengxian Ji, Jingpu Yang, Zirui Song +6
With the rapid advancement of generative artificial intelligence technology, Graphical User Interface (GUI) agents have demonstrated tremendous potential for autonomously managing…
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
Zirui Song, Guangxian Ouyang, Mingzhe Li +10
Large Vision-Language Models (LVLMs) have recently advanced robotic manipulation by leveraging vision for scene perception and language for instruction following. However, existing…