works on

From the 1 of 14 linked papers with an AI index.

collaborators

14 papers

cs.CV2026

Flow Matching in Feature Space for Stochastic World Modeling

Francois Porcher, Nicolas Carion, Karteek Alahari +1

The paper introduces FlowWM, a stochastic world model that applies flow matching directly in high‑dimensional pretrained feature spaces (e.g., DINOv3) to improve future forecasting…

cs.RO2026

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

Shizhe Chen, Paul Pacaud, Cordelia Schmid

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most exi…

cs.CV2026

MAGICIAN: Efficient Long-Term Planning with Imagined Gaussians for Active Mapping

Shiyao Li, Antoine Guédon, Shizhe Chen +1

Active mapping aims to determine how an agent should move to efficiently reconstruct unknown environments. Most existing approaches rely on greedy next-best-view prediction, result…

cs.CV2026

HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching

Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen +3

Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plau…

cs.RO2026

Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation

Paul Pacaud, Ricardo Garcia, Shizhe Chen +1

Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in robot failure detection, their generaliz…

cs.SD2025

Hear: Hierarchically Enhanced Aesthetic Representations For Multidimensional Music Evaluation

Shuyang Liu, Yuan Jin, Rui Lin +3

Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evalu…