55 papers
EgoPlay: Event-Triggered Video Editing for Egocentric Streams
Jinjie Mai, Gordon Guocheng Qian, Willi Menapace +8
We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data buil…
TG-Diff: Coupling Discrete Topology Diffusion and Topology-conditioned Geometry Diffusions for B-Rep Generation
MingZe Sun, Haiyong Jiang, Bingchen Yang +4
Boundary representation (B-rep) is the standard format for computer-aided design (CAD). This article proposes a lightweight two-stage diffusion-based B-rep generation framework, TG…
RegHead: Non-Humanoid Head Blendshapes via Feed-Forward Registration
Jiahao Luo, Hao Zhang, Jianqi Chen +9
RegHead is a framework that builds semantic blendshape sets for animatable non‑humanoid head avatars using a fast feed‑forward registration model and a large dataset of shared expr…
CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging
Zhenyu Li, Aleksandar Cvejic, Zehui Chen +1
Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond functional correctness. While…
RAGA: Real Time Ray Traced Gaussian Shadow Casting for 3DGS Avatar-Scene Interaction
Aymen Mir, Riza Alp Guler, Jian Wang +3
We study the problem of physically plausible shadow casting when animating 3D Gaussian Splatting (3DGS) avatars, either individually or in multi-avatar and object-interaction scena…
AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors
Aymen Mir, Riza Alp Guler, Xiangjun Tang +2
We present AHOY, a method for reconstructing complete, animatable 3D Gaussian avatars from in-the-wild monocular video despite heavy occlusion. Existing methods assume unoccluded i…