3 papers
cs.CV2025
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
Qihang Ma, Shengyu Li, Jie Tang +5
Multi-modal keyphrase prediction (MMKP) aims to advance beyond text-only methods by incorporating multiple modalities of input information to produce a set of conclusive phrases. T…
cs.RO2024
iKalibr-RGBD: Partially-Specialized Target-Free Visual-Inertial Spatiotemporal Calibration For RGBDs via Continuous-Time Velocity Estimation
Shuolong Chen, Xingxing Li, Shengyu Li +1
Visual-inertial systems have been widely studied and applied in the last two decades (from the early 2000s to the present), mainly due to their low cost and power consumption, smal…
cs.RO2024
DBA-Fusion: Tightly Integrating Deep Dense Visual Bundle Adjustment with Multiple Sensors for Large-Scale Localization and Mapping
Yuxuan Zhou, Xingxing Li, Shengyu Li +3
Visual simultaneous localization and mapping (VSLAM) has broad applications, with state-of-the-art methods leveraging deep neural networks for better robustness and applicability.…