activity
20232025
most citedMulti-Transmotion: Pre-trained Model for Human Motion Prediction

2 citations · 6 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2025

Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

Wuyang Li, Wentao Pan, Po-Chien Luan +2

We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with high temporal consistency, plausible scene transitions, and controllable streaming story…

cs.CV2025

OmniTraj: Pre-Training on Heterogeneous Data for Adaptive and Zero-Shot Human Trajectory Prediction

Yang Gao, Po-Chien Luan, Kaouther Messaoud +2

While large-scale pre-training has advanced human trajectory prediction, a critical challenge remains: zero-shot transfer to unseen dataset with varying temporal dynamics. State-of…

cs.CV2025

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models

Pablo Acuaviva, Aram Davtyan, Mariam Hassan +4

Video Diffusion Models (VDMs) have emerged as powerful generative tools, capable of synthesizing high-quality spatiotemporal content. Yet, their potential goes far beyond mere vide…

cs.CV2025

COARSE: Collaborative Pseudo-Labeling with Coarse Real Labels for Off-Road Semantic Segmentation

Aurelio Noca, Xianmei Lei, Jonathan Becktor +6

Autonomous off-road navigation faces challenges due to diverse, unstructured environments, requiring robust perception with both geometric and semantic understanding. However, scar…

cs.CV2025

Unified Human Localization and Trajectory Prediction with Monocular Vision

Po-Chien Luan, Yang Gao, Celine Demonsant +1

Conventional human trajectory prediction models rely on clean curated data, requiring specialized equipment or manual labeling, which is often impractical for robotic applications.…

cs.CV2024

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Mariam Hassan, Sebastian Stapf, Ahmad Rahimi +17

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, ou…