4 papers
Robotic Manipulation is Vision-to-Geometry Mapping: Vision-Geometry Backbones over Language and Video Models
Zijian Song, Qichang Li, Jiawei Zhou +4
At its core, robotic manipulation is a problem of vision-to-geometry mapping (). Physical actions are fundamentally defined by geometric properties like 3D posi…
Learning Hierarchical and Geometry-Aware Graph Representations for Text-to-CAD
Shengjie Gong, Wenjie Peng, Hongyuan Chen +5
Text-to-CAD code generation is a long-horizon task that translates textual instructions into long sequences of interdependent operations. Existing methods typically decode text dir…
ReplayCAD: Generative Diffusion Replay for Continual Anomaly Detection
Lei Hu, Zhiyong Gan, Ling Deng +4
Continual Anomaly Detection (CAD) enables anomaly detection models in learning new classes while preserving knowledge of historical classes. CAD faces two key challenges: catastrop…
Monocular and Generalizable Gaussian Talking Head Animation
Shengjie Gong, Haojie Li, Jiapeng Tang +5
In this work, we introduce Monocular and Generalizable Gaussian Talking Head Animation (MGGTalk), which requires monocular datasets and generalizes to unseen identities without per…