works on

From the 7 of 78 linked papers with an AI index.

collaborators

78 papers

cs.CV2026

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

Wanshun Su, Yang Shi, Feihu Liu +10

Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio tok…

cs.CV2026

CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation

Yizhuo Jia, Jingyun Hua, Yuanxing Zhang

Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prio…

cs.CV2026

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Qixun Wang, Yang Shi, Letian Cheng +11

The paper introduces Beacon, an agentic visual reasoning system that learns when to invoke external tools and how to use them effectively, improving multimodal large language model…

cs.CV2026

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Tengfei Liu, Yang Shi, Yuran Wang +16

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-…

cs.LG2026

Flux-OPD: On-Policy Distillation with Evolving Contexts

Yuran Wang, Zekun Wang, Bohan Zeng +10

The paper introduces Flux-OPD, a method for training large language models by distilling knowledge from teachers while using evolving contexts as supervision, and stabilizes the pr…

cs.CV2026

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

Yifan Xu, Zihao Wang, Zhixiao Wang +6

Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occ…