6 papers
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
Seong Jong Yoo, Siyuan Peng, Felix Gu +2
Choreographic motion generation poses unique challenges for AI, demanding precise semantic control over complex, temporally structured, and expressive full-body dynamics. While exi…
Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
Jiayi Wu, Haoming Cai, Cornelia Fermuller +2
While Video Diffusion Models (VDMs) excel at synthesizing high-fidelity videos, enabling precise camera and scene control remains challenging. Existing methods predominantly rely o…
Automated Detection of Multiple Sclerosis Lesions on 7-tesla MRI Using U-net and Transformer-based Segmentation
Michael Maynord, Minghui Liu, Cornelia Fermüller +4
Ultra-high field 7-tesla (7T) MRI improves visualization of multiple sclerosis (MS) white matter lesions (WML) but differs sufficiently in contrast and artifacts from 1.5-3T imagin…
From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
Jingxi Chen, Yixiao Zhang, Xiaoye Qian +4
Images can be viewed as layered compositions, foreground objects over background, with potential occlusions. This layered representation enables independent editing of elements, of…
NatSGLD: A Dataset with Speech, Gesture, Logic, and Demonstration for Robot Learning in Natural Human-Robot Interaction
Snehesh Shrestha, Yantian Zha, Saketh Banagiri +3
Recent advances in multimodal Human-Robot Interaction (HRI) datasets emphasize the integration of speech and gestures, allowing robots to absorb explicit knowledge and tacit unders…
VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference
Seong Jong Yoo, Snehesh Shrestha, Irina Muresanu +1
Musicians delicately control their bodies to generate music. Sometimes, their motions are too subtle to be captured by the human eye. To analyze how they move to produce the music,…