collaborators

11 papers

cs.CE2026

MMDynOpt-Agent: Dynamic Optimization for Multimodal Large Language Model Reasoning via Reinforcement Learning

Wenjin Liu, Haoran Luo, Fayuan Ke +5

Recently, multimodal large language models (MLLMs) have demonstrated strong potential in visual understanding and complex reasoning tasks. However, existing methods often struggle…

cs.CV2026

Gaussian-Voxel Duet: A Dual-Scaffolding Hybrid Representation for Fast and Accurate Monocular Surface Reconstruction

Zhenhua Du, Zhen Tan, Haoyu Zhang +3

While 3D Gaussian Splatting has achieved remarkable success in photorealistic novel view synthesis, its pursuit of fast and high-fidelity 3D reconstruction has long been constraine…

cs.CV2026

ConFi-GS Confidence-Guided High-Frequency Injection for 3D Gaussian Splatting Super-Resolution

Jiaxiang Li, Zongtan Zhou, Zhen Tan +2

Reconstructing high-quality 3D scenes from low-resolution multi-view images remains challenging for 3D Gaussian Splatting (3DGS), because insufficient high-frequency observations o…

cs.RO2026

Event-based SLAM Benchmark for High-Speed Maneuvers

Sheng Zhong, Junkai Niu, Guillermo Gallego +7

Event-based cameras are bio-inspired sensors with pixels that independently and asynchronously respond to brightness changes at microsecond resolution, offering the potential to ha…

cs.CV2026

GESS: Multi-cue Guided Local Feature Learning via Geometric and Semantic Synergy

Yang Yi, Xieyuanli Chen, Jinpu Zhang +2

Robust local feature detection and description are foundational tasks in computer vision. Existing methods primarily rely on single appearance cues for modeling, leading to unstabl…

cs.CV2026

TAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events

Jiaxiong Liu, Zhen Tan, Jinpu Zhang +4

Tracking any point (TAP) is a fundamental yet challenging task in computer vision, requiring high precision and long-term motion reasoning. Recent attempts to combine RGB frames an…