works on

From the 2 of 8 linked papers with an AI index.

collaborators

8 papers

cs.CV2026

AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

Yanghai Wang, Jiahao Wang, Jiafu Tang +9

The paper introduces AVSCap, a system for omni-modal video captioning that explicitly binds visual and audio events, using a large tri-modal dataset and a two-stage training with r…

cs.CY2026

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs

Anqi Li, Jie Zhang, Zhongqi Wang +4

While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predomi…

cs.AI2026

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Hongcheng Gao, Hailong Qu, Jingyi Tang +18

Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predomin…

cs.CV2026

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

Jiahao Wang, An Ping, Yanghai Wang +13

While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…

cs.CV2026

A Survey of Multimodal Hallucination Evaluation and Detection

Zhiyuan Chen, Yuecong Min, Jie Zhang +4

Multi-modal Large Language Models (MLLMs) have emerged as a powerful paradigm for integrating visual and textual information, supporting a wide range of multi-modal tasks. However,…

cs.CL2025

JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation

Fan Xu, Huixuan Zhang, Zhenliang Zhang +2

Current large language models (LLMs) often suffer from hallucination issues, i,e, generating content that appears factual but is actually unreliable. A typical hallucination detect…