collaborators

5 papers

cs.CV2026

InfraNet: Quality-Aware RGB Guidance for Efficient Infrared Object Detection

Zichao Feng, Haodong Zhu, Jingying Yang +8

Robust object detection under adverse visual conditions remains a long-standing challenge for multi-modal perception systems. Existing fusion-based methods typically require both R…

cs.CV2026

Teaching Vision-Language-Action Models What to See and Where to Look

Yuguang Yang, Canyu Chen, Zhewen Tan +10

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual q…

cs.GR2025

Surf3R: Rapid Surface Reconstruction from Sparse RGB Views in Seconds

Haodong Zhu, Changbai Li, Yangyang Ren +5

Current multi-view 3D reconstruction methods rely on accurate camera calibration and pose estimation, requiring complex and time-intensive pre-processing that hinders their practic…

cs.CV2025

WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection

Haodong Zhu, Wenhao Dong, Linlin Yang +10

Leveraging the complementary characteristics of visible (RGB) and infrared (IR) imagery offers significant potential for improving object detection. In this paper, we propose WaveM…

cs.LG2025

Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method

Qingcheng Zhu, Yangyang Ren, Linlin Yang +9

Deploying large language models (LLMs) is challenging due to their massive parameters and high computational costs. Ultra low-bit quantization can significantly reduce storage and…