activity
20242026
collaborators

9 papers

cs.CV2026

CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

Jiaye Fu, Weiqi Li, Qiankun Gao +5

Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and g…

cs.CV2026

CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images

Haijie Li, Jiaxin Zhang, Dave Zhenyu Chen +3

Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods…

cs.CV2026

Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction

Zecheng Tang, Jiaye Fu, Qiankun Gao +5

Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input images, but scaling them to video…

eess.IV2025

ReCon-GS: Continuum-Preserved Gaussian Streaming for Fast and Compact Reconstruction of Dynamic Scenes

Jiaye Fu, Qiankun Gao, Chengxiang Wen +4

Online free-viewpoint video (FVV) reconstruction is challenged by slow per-frame optimization, inconsistent motion estimation, and unsustainable storage demands. To address these c…

cs.CV2025

InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception

Haijie Li, Yanmin Wu, Jiarui Meng +4

3D scene understanding has become an essential area of research with applications in autonomous driving, robotics, and augmented reality. Recently, 3D Gaussian Splatting (3DGS) has…

cs.CV2024

HiCoM: Hierarchical Coherent Motion for Streamable Dynamic Scene with 3D Gaussian Splatting

Qiankun Gao, Jiarui Meng, Chengxiang Wen +2

The online reconstruction of dynamic scenes from multi-view streaming videos faces significant challenges in training, rendering and storage efficiency. Harnessing superior learnin…