4 papers
Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment
Pengxu Chen, Yao Zhu, Guangming Zhu +4
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding. However, they remain prone to hallucinations, generating responses that…
ReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D Reconstruction
Xinze Li, Yiyuan Wang, Pengxu Chen +4
Streaming 3D reconstruction relies on a compact recurrent scene state to process long image streams in linear time and bounded memory. However, repeated updates can gradually corru…
QuadBox: Accelerating 3D Gaussian Splatting with Geometry-Aware Boxes
Xinze Li, Bohan Yang, Pengxu Chen +4
3D Gaussian Splatting (3DGS) has emerged as an advanced technique for real-time novel view synthesis by representing scene geometry and appearance using differentiable Gaussian pri…
S-VGGT: Structure-Aware Subscene Decomposition for Scalable 3D Foundation Models
Xinze Li, Pengxu Chen, Yiyuan Wang +2
Feed-forward 3D foundation models face a key challenge: the quadratic computational cost introduced by global attention, which severely limits scalability as input length increases…