Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
InTraGen: Trajectory-controlled Video Generation for Object Interactions
Zuhao Liu, Aleksandar Yanev, Ahmad Mahmood +7
Advances in video generation have significantly improved the realism and quality of created scenes. This has fueled interest in developing intuitive tools that let users leverage v…
cs.CV2025
Dynamic Object Masks as Goal Representations for Visual Goal-Conditioned Reinforcement Learning
Fahim Shahriar, Cheryl Wang, Alireza Azimi +6
Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-based goal representations that ar…
cs.CV2025
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
Ahmad Mahmood, Ashmal Vayani, Muzammal Naseer +2
Recent studies have demonstrated the effectiveness of Large Language Models (LLMs) as reasoning modules that can deconstruct complex tasks into more manageable sub-tasks, particula…