collaborators

12 papers

cs.CV2026

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

Penghui Yang, Long Xing, Xiaoyi Dong +10

Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Curren…

cs.CV2026

Harnessing Streaming Video in the Wild

Dingyu Yao, Shuhuan Gu, Qingyi Si +8

Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An i…

cs.CV2026

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

Jiahao Wang, An Ping, Yanghai Wang +13

While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…

cs.CV2026

Light-WAM: Efficient World Action Models with State-Fusion Action Decoding

Ziang Li, Dongzhou Cheng, Yibin Wang +5

World Action Models (WAMs) extend robot policy learning by incorporating future prediction as an additional training objective, encouraging the policy to encode task-relevant tempo…

cs.CV2026

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO

Jiazi Bu, Pengyang Ling, Yujie Zhou +8

Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. However, we have identified that t…

cs.CV2026

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion

Feng Han, Zhixiong Zhang, Zheming Liang +2

Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image-text training aimed at multim…