12 papers
SR-LIO++: LiDAR-Inertial Odometry and Quantized Mapping with Caching-Aware Sweep Reconstruction
Zikang Yuan, Ruiye Ming, Chengwei Zhao +6
Addressing the inherent low acquisition frequency limitation of 3D LiDAR to achieve high-frequency output has become a critical research focus in the LiDAR-Inertial Odometry (LIO)…
Path-Decoupled Hyperbolic Flow Matching for Few-Shot Adaptation
Lin Li, Ziqi Jiang, Gefan Ye +5
Recent advances in cross-modal few-shot adaptation treat visual-semantic alignment as a continuous feature transport problem via Flow Matching (FM). However, we argue that Euclidea…
A Deployment-Friendly Foundational Framework for Efficient Computational Pathology
Yu Cai, Cheng Jin, Jiabo Ma +25
Pathology foundation models (PFMs) generalize well across computational pathology tasks but remain costly for gigapixel whole-slide image analysis. Here, we present LitePath, a dep…
A Study of Finetuning Video Transformers for Multi-view Geometry Tasks
Huimin Wu, Kwang-Ting Cheng, Stephen Lin +1
This paper presents an investigation of vision transformer learning for multi-view geometry tasks, such as optical flow estimation, by fine-tuning video foundation models. Unlike p…
Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension
Lin Li, Wei Chen, Jiahui Li +2
Recent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relati…
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
Xixi Jiang, Chen Yang, Dong Zhang +3
Vision Transformer models have shown impressive effectiveness in the surgical video understanding tasks through long-range dependency modeling. However, current methods suffer from…