4 papers
SEAM: Smooth Execution of Action-Chunked Motion for Vision-Language-Action Policies
Dijia Zhan, Xuemiao Xu, Jinyi Li +1
Vision-Language-Action (VLA) policies that execute fixed-length action chunks can exhibit multimodal bifurcation: a cross-chunk inconsistency in which adjacent chunks generated fro…
NavOne: One-Step Global Planning for Vision-Language Navigation on Top-Down Maps
Dijia Zhan, Jinyi Li, Chenxi Zheng +4
Existing Vision-Language Navigation (VLN) methods typically adopt an egocentric, step-by-step paradigm, which struggles with error accumulation and limits efficiency. While recent…
HyperTea: A Hypergraph-based Temporal Enhancement and Alignment Network for Moving Infrared Small Target Detection
Zhaoyuan Qi, Weihua Gao, Wenlong Niu +3
In practical application scenarios, moving infrared small target detection (MIRSTD) remains highly challenging due to the target's small size, weak intensity, and complex motion pa…
Point-to-Mask: From Arbitrary Point Annotations to Mask-Level Infrared Small Target Detection
Weihua Gao, Wenlong Niu, Jie Tang +3
Infrared small target detection (IRSTD) methods predominantly formulate the task as pixel-level segmentation, which requires costly dense annotations and is not well suited to tiny…