5 papers
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
Xuanzhao Dong, Wenhui Zhu, Xiwen Chen +13
The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to support clinical diagnosis. However,…
DeepSVU: Towards In-depth Security-oriented Video Understanding via Unified Physical-world Regularized MoE
Yujie Jin, Wenxin Zhang, Jingjing Wang +1
In the literature, prior research on Security-oriented Video Understanding (SVU) has predominantly focused on detecting and localize the threats (e.g., shootings, robberies) in vid…
Towards LLM-centric Affective Visual Customization via Efficient and Precise Emotion Manipulating
Jiamin Luo, Xuqian Gu, Jingjing Wang +1
Previous studies on visual customization primarily rely on the objective alignment between various control signals (e.g., language, layout and canny) and the edited images, which l…
Table-Critic: A Multi-Agent Framework for Collaborative Criticism and Refinement in Table Reasoning
Peiying Yu, Guoxin Chen, Jingjing Wang
Despite the remarkable capabilities of large language models (LLMs) in various reasoning tasks, they still struggle with table reasoning tasks, particularly in maintaining consiste…
Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLM
Junxiao Ma, Jingjing Wang, Jiamin Luo +2
Prior studies on Video Anomaly Detection (VAD) mainly focus on detecting whether each video frame is abnormal or not in the video, which largely ignore the structured video semanti…