From the 1 of 5 linked papers with an AI index.
5 papers
AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning
Yanghai Wang, Jiahao Wang, Jiafu Tang +9
The paper introduces AVSCap, a system for omni-modal video captioning that explicitly binds visual and audio events, using a large tri-modal dataset and a two-stage training with r…
OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Reliability Calibration for Long-Form Omni Hallucination
Zixuan Dong, Jiafu Tang, Zhide Lei +7
Long-video Omni assistants often fail not by inventing content, but by misbinding real evidence: they hear the right utterance and see the right event, yet attach it to the wrong s…
QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents
Ye Yuan, Rui Song, Weien Li +12
Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environ…
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
Zixuan Dong, Baoyun Peng, Yufei Wang +4
Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video…
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
Xinxin Dong, Baoyun Peng, Haokai Ma +4
Video Question Answering (VideoQA) requires identifying sparse critical moments in long videos and reasoning about their causal relationships to answer semantically complex questio…