Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum
Xingjian Wang, Shijian Wang, Yibo Wang +4
Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely ov…
cs.CV2025
Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
Feng Wang, Zihao Yu
Reinforcement Learning (RL) has recently emerged as a powerful technique for improving image and video generation in Diffusion and Flow Matching models, specifically for enhancing…