works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CV2026

AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

Yanghai Wang, Jiahao Wang, Jiafu Tang +9

The paper introduces AVSCap, a system for omni-modal video captioning that explicitly binds visual and audio events, using a large tri-modal dataset and a two-stage training with r…

cs.MM2026

OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Reliability Calibration for Long-Form Omni Hallucination

Zixuan Dong, Jiafu Tang, Zhide Lei +7

Long-video Omni assistants often fail not by inventing content, but by misbinding real evidence: they hear the right utterance and see the right event, yet attach it to the wrong s…

cs.CL2026

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

Ye Yuan, Rui Song, Weien Li +12

Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environ…

cs.CV2025

See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops

Zixuan Dong, Baoyun Peng, Yufei Wang +4

Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video…

cs.CV2025

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering

Xinxin Dong, Baoyun Peng, Haokai Ma +4

Video Question Answering (VideoQA) requires identifying sparse critical moments in long videos and reasoning about their causal relationships to answer semantically complex questio…