6 citations · 16 across the 39 of their papers we have counts for
36 papers · 1 filter
Technical Report on the CVPR 2026@AdvML Workshop Challenge
Tianyuan Zhang, Zonglei Jing, Jiangfan Liu +47
Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Cha…
RATS! Patches Talk Through Registers: Emergent Parts in Register Attention Transformers
Timing Yang, Predrag Neskovic, Jansen Seheult +4
When humans see a bird, they recognize far more than just "bird" -- they see a head, wings, and talons, a structured assembly of reusable parts that can be identified across every…
SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images
Sky Cen, Wufei Ma, Guofeng Zhang +2
Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it remains unclear whether the…
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
Soumava Paul, Prakhar Kaushik, Alan Yuille
Multiview 3D evaluation assumes that the images being scored are observations of one static 3D scene. This assumption can fail in NVS and sparse-view reconstruction: inputs or gene…
HECTOR: Hybrid Editable Compositional Object References for Video Generation
Guofeng Zhang, Angtian Wang, Jacob Zhiyuan Fang +4
Real-world videos naturally portray complex interactions among distinct physical objects, effectively forming dynamic compositions of visual elements. However, most current video g…
Thinking with Spatial Code for Physical-World Video Reasoning
Jieneng Chen, Wenxin Ma, Ruisheng Yuan +3
We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. W…