From the 1 of 14 linked papers with an AI index.
14 papers
Flow Matching in Feature Space for Stochastic World Modeling
Francois Porcher, Nicolas Carion, Karteek Alahari +1
The paper introduces FlowWM, a stochastic world model that applies flow matching directly in high‑dimensional pretrained feature spaces (e.g., DINOv3) to improve future forecasting…
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
Shizhe Chen, Paul Pacaud, Cordelia Schmid
Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most exi…
MAGICIAN: Efficient Long-Term Planning with Imagined Gaussians for Active Mapping
Shiyao Li, Antoine Guédon, Shizhe Chen +1
Active mapping aims to determine how an agent should move to efficiently reconstruct unknown environments. Most existing approaches rely on greedy next-best-view prediction, result…
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen +3
Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plau…
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
Paul Pacaud, Ricardo Garcia, Shizhe Chen +1
Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in robot failure detection, their generaliz…
Hear: Hierarchically Enhanced Aesthetic Representations For Multidimensional Music Evaluation
Shuyang Liu, Yuan Jin, Rui Lin +3
Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evalu…