works on

From the 1 of 11 linked papers with an AI index.

collaborators

11 papers

cs.CV2026

AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

Yanghai Wang, Jiahao Wang, Jiafu Tang +9

The paper introduces AVSCap, a system for omni-modal video captioning that explicitly binds visual and audio events, using a large tri-modal dataset and a two-stage training with r…

cs.CV2026

STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Model

Hailan Zhang, Haipeng Liu, Bo Fu +1

Although pretrained text-to-image (T2I) generation models can produce high-quality images, they often fail to faithfully reflect the semantic intent of complex prompts due to stoch…

cs.CV2026

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

Jiahao Wang, An Ping, Yanghai Wang +13

While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…

cs.CV2026

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

Lin Fu, Zheyuan Yang, Yang Wang +3

We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding. It comprises 315K video reas…

cs.CV2026

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Zhe Cao, Tao Wang, Jiaming Wang +10

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented,…

cs.MM2026

OmniHalluc-L: Counterfactual Benchmarking and Modality-Perturbation Reliability Calibration for Long-Form Omni Hallucination

Zixuan Dong, Jiafu Tang, Zhide Lei +7

Long-video Omni assistants often fail not by inventing content, but by misbinding real evidence: they hear the right utterance and see the right event, yet attach it to the wrong s…