8 papers
STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning
Zinuo Li, Yongxin Guo, Jun Liu +7
Human understanding of video dynamics relies on forming structured representations of entities, actions, and temporal relations before engaging in abstract reasoning. In contrast,…
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
Ashkan Taghipour, Morteza Ghahremani, Zinuo Li +3
Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally…
EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines
Shuo Zhang, Chaofa Yuan, Ryan Guo +11
While LLM-based agents have shown promise for deep research, most existing approaches rely on fixed workflows that struggle to adapt to real-world, open-ended queries. Recent work…
Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
Zinuo Li, Xian Zhang, Yongxin Guo +5
Humans naturally understand moments in a video by integrating visual and auditory cues. For example, localizing a scene in the video like "A scientist passionately speaks on wildli…
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
Xian Zhang, Zexi Wu, Zinuo Li +5
Understanding long-form videos remains a significant challenge for vision--language models (VLMs) due to their extensive temporal length and high information density. Most current…
LatentMove: Towards Complex Human Movement Video Generation
Ashkan Taghipour, Morteza Ghahremani, Mohammed Bennamoun +5
Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often s…