2 papers
cs.CV2026
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
Peiran Wu, Zhuorui Yu, Yunze Liu +3
The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. However, visual language models (VLMs) still face heavy computational costs when e…
cs.CV2025
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
Peiran Wu, Yunze Liu, Zhengdong Zhu +2
Real-world user-generated videos, especially on platforms like TikTok, often feature rich and intertwined audio visual content. However, existing video captioning benchmarks and mo…