1 citations · 1 across the 11 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Video Models Can Reason with Verifiable Rewards
Tinghui Zhu, Sheng Zhang, James Y. Huang +5
Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable re…
cs.CV2026
When Vision Speaks for Sound
Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6
Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate…