collaborators

7 papers

cs.CL2026

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen +4

Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies…

cs.RO2026

Graph-Loc: Robust Graph-Based LiDAR Pose Tracking with Compact Structural Map Priors under Low Observability and Occlusion

Wentao Zhao, Yihe Niu, Zikun Chen +4

Map-based LiDAR pose tracking is essential for long-term autonomous operation, where onboard map priors need be compact for scalable storage and fast retrieval, while online observ…

cs.CV2026

OTPL-VIO: Robust Visual-Inertial Odometry with Optimal Transport Line Association and Adaptive Uncertainty

Zikun Chen, Wentao Zhao, Yihe Niu +2

Robust stereo visual-inertial odometry (VIO) remains challenging in low-texture scenes and under abrupt illumination changes, where point features become sparse and unstable, leadi…

cs.CV2025

Guided Diffusion-based Generation of Adversarial Objects for Real-World Monocular Depth Estimation Attacks

Yongtao Chen, Yanbo Wang, Wentao Zhao +3

Monocular Depth Estimation (MDE) serves as a core perception module in autonomous driving systems, but it remains highly susceptible to adversarial attacks. Errors in depth estimat…

cs.CV2025

MUT3R: Motion-aware Updating Transformer for Dynamic 3D Reconstruction

Guole Shen, Tianchen Deng, Xingrui Qin +6

Recent stateful recurrent neural networks have achieved remarkable progress on static 3D reconstruction but remain vulnerable to motion-induced artifacts, where non-rigid regions c…

cs.CV2025

GRS-SLAM3R: Real-Time Dense SLAM with Gated Recurrent State

Guole Shen, Tianchen Deng, Yanbo Wang +4

DUSt3R-based end-to-end scene reconstruction has recently shown promising results in dense visual SLAM. However, most existing methods only use image pairs to estimate pointmaps, o…