1 citations · 1 across the 4 of their papers we have counts for
4 papers
StreamFlow: Dynamic Memory Flows for Streaming Video Understanding
Muxin Fu, Yifan Zhang, Wentao Zhang +5
Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under strict causality and bounded…
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Jiaxing Li, Kai Zou, Cindy Zhou +7
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initia…
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
Zhiyuan Hu, Shiyun Xiong, Yifan Zhang +5
Recent advancements in visual language models (VLMs) have notably enhanced their capabilities in handling complex Graphical User Interface (GUI) interaction tasks. Despite these im…
MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation
Longtao Zheng, Yifan Zhang, Hanzhong Guo +6
Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, ma…