7 papers
MechVerse: Evaluating Physical Motion Consistency in Video Generation Models
Rahul Jain, Mayank Patel, Asim Unmesh +1
Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion governed by kinematic and ge…
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
Asim Unmesh, Kaki Ramesh, Mayank Patel +2
Temporal Action Segmentation (TAS) requires dividing videos into action segments, yet the vast space of activities and alternative breakdowns makes collecting comprehensive dataset…
DYNAMO: Dependency-Aware Deep Learning Framework for Articulated Assembly Motion Prediction
Mayank Patel, Rahul Jain, Asim Unmesh +1
Understanding the motion of articulated mechanical assemblies from static geometry remains a core challenge in 3D perception and design automation. Prior work on everyday articulat…
Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
Runlin Duan, Yuzhao Chen, Rahul Jain +3
Generative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including ob…
An Exploratory Study on Multi-modal Generative AI in AR Storytelling
Hyungjun Doh, Jingyu Shi, Rahul Jain +2
Storytelling in AR has gained attention due to its multi-modality and interactivity. However, generating multi-modal content for AR storytelling requires expertise and efforts for…
Visualizing Causality in Mixed Reality for Manual Task Learning: An Exploratory Study
Rahul Jain, Jingyu Shi, Andrew Benton +4
Mixed Reality (MR) is gaining prominence in manual task skill learning due to its in-situ, embodied, and immersive experience. To teach manual tasks, current methodologies break th…