10 papers
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
Yunheng Li, Guohong Mu, Hao Li +4
Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on large multi-task datasets remains challenging, as…
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Xumin Yu, Zuyan Liu, Zhenyu Yang +5
A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete s…
LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams
Zhenyu Yang, Kairui Zhang, Bing Wang +2
Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide…
Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets
Zhenyu Yang, Zemin Du, Shengsheng Qian +1
Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption without training samples. Existing…
Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding
Zhenyu Yang, Kairui Zhang, Shengsheng Qian +2
Online Video Large Language Models (Video-LLMs) have advanced toward seamless human-AI interaction through frame-by-frame processing and proactive responding. However, a critical c…
SoMe: A Realistic Benchmark for LLM-based Social Media Agents
Dizhan Xue, Jing Cui, Shengsheng Qian +2
Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM…