5 papers · 1 filter
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
Dingyu Yao, Junhao Zhou, Chenxu Yang +12
Many moments in the real world do not wait for a user to ask. A fire starts on a security monitor, an expression flickers across a video call, or a product a viewer wants flashes b…
Harnessing Streaming Video in the Wild
Dingyu Yao, Shuhuan Gu, Qingyi Si +8
Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An i…
Online Self-Calibration Against Hallucination in Vision-Language Models
Minghui Chen, Chenxu Yang, Hengjie Zhu +3
Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image. Recent preference alignment…
EasyVideoR1: Easier RL for Video Understanding
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
Reinforcement learning from verifiable rewards (RLVR) has demonstrated remarkable effectiveness in improving the reasoning capabilities of large language models. As models evolve i…
Channel Attention-Guided Cross-Modal Knowledge Distillation for Referring Image Segmentation
Chen Yang
Referring image segmentation (RIS) requires accurate segmentation of target regions in images according to language descriptions, which is a cross-modal task integrating vision and…