collaborators

11 papers

cs.CV2026

ReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D Reconstruction

Xinze Li, Yiyuan Wang, Pengxu Chen +4

Streaming 3D reconstruction relies on a compact recurrent scene state to process long image streams in linear time and bounded memory. However, repeated updates can gradually corru…

cs.AI2026

PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft

Yuchen Guo, Junli Gong, Weicheng Wang +3

We present PEAM, a Parametric Embodied Agent Memory framework in Minecraft that transforms agent memory from inference-time retrieval into parameter-resident skills internalized th…

cs.CV2026

Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought

Yuchen Guo, Junli Gong, Hongmin Cai +2

Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instructions remain target-referentia…

cs.CV2026

Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment

Yuchen Guo, Junli Gong, Yao Lu +3

Infrared-Visible image fusion (IVIF) aims to integrate thermal information and detailed spatial structures into a single fused image to enhance perception. However, existing evalua…

cs.CV2026

Adding Thermal Awareness to Visual Systems in Real-Time via Distilled Diffusion Models

Yuchen Guo, Junli Gong, Wenjun Dong +2

Purely RGB-based vision models often fail to provide reliable cues in challenging scenarios such as nighttime and fog, leading to degraded performance and safety risks. Infrared im…

cs.CV2026

QuadBox: Accelerating 3D Gaussian Splatting with Geometry-Aware Boxes

Xinze Li, Bohan Yang, Pengxu Chen +4

3D Gaussian Splatting (3DGS) has emerged as an advanced technique for real-time novel view synthesis by representing scene geometry and appearance using differentiable Gaussian pri…