9 papers
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Qianggang Ding, Xingyao Wang, Rui Feng +20
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person fo…
Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining
Guanghao Zhu, Zeyu Liu, Zhitian Hou +10
Biomedical figures are explained not by captions alone but by body-text passages that discuss them. Yet current multimodal corpora typically reduce figures to isolated image-captio…
TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
Minheng Ni, Zhengyuan Yang, Yaowen Zhang +9
We study technical image generation, where a model must synthesize information-dense, scientifically precise illustrations from detailed descriptions rather than merely produce vis…
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
Minheng Ni, Yutao Fan, Zhengyuan Yang +6
Recent advances in large multimodal models (LMMs) have enabled instruction-based image editing, allowing users to modify visual content via natural language descriptions. However,…
ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models
Lei Zhang, Ju Dong, Kaixin Bai +4
Recent advances in large multimodal models have enabled new opportunities in embodied AI, particularly in robotic manipulation. These models have shown strong potential in generali…
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
Bowen Dong, Minheng Ni, Zitong Huang +3
Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse…