Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
Siyoon Jin, Seongchan Kim, Dahyun Chung +5
Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internall…
cs.CV2025
Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
Junsu Kim, Naeun Kim, Jaeho Lee +3
The reasoning-based pose estimation (RPE) benchmark has emerged as a widely adopted evaluation standard for pose-aware multimodal large language models (MLLMs). Despite its signifi…