From the 1 of 6 linked papers with an AI index.
6 papers
SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
Yunzhan Fu, Enyu Bao, Xiangyu Shen +4
The paper introduces SCALPEL, a framework that fine‑tunes a generative medical LLM into an isotropic encoder and uses an asymmetric, anatomy‑negation‑aware contrastive objective to…
Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
Yihao Wu, Chenyi Xu, Liqi Yan +6
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multim…
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
He Zhang, Lingzhu Xiang, Haitao Lin +23
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pr…
LongDPM: Overlap-Aware 4D Reconstruction from Long Monocular Videos
Chenyi Xu, Yihao Wu, Liqi Yan +4
Recovering a dynamic 3D scene from a long monocular video is crucial for dense geometry, camera motion, and temporal correspondence to remain consistent in a shared coordinate syst…
MetaWild: A Multimodal Dataset for Animal Re-Identification with Environmental Metadata
Yuzhuo Li, Di Zhao, Tingrui Qiao +3
Identifying individual animals within large wildlife populations is essential for effective wildlife monitoring and conservation efforts. Recent advancements in computer vision hav…
An Individual Identity-Driven Framework for Animal Re-Identification
Yihao Wu, Di Zhao, Jingfeng Zhang +1
Reliable re-identification of individuals within large wildlife populations is crucial for biological studies, ecological research, and wildlife conservation. Classic computer visi…