1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2026
VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation
Yijie Zhu, Jie He, Rui Shao +4
Recent vision-language-action (VLA) models have significantly advanced robotic manipulation by unifying perception, reasoning, and control. To achieve such integration, recent stud…
cs.CV2025
AdaMHF: Adaptive Multimodal Hierarchical Fusion for Survival Prediction
Shuaiyu Zhang, Xun Lin, Rongxiang Zhang +5
The integration of pathologic images and genomic data for survival analysis has gained increasing attention with advances in multimodal learning. However, current methods often ign…
cs.CV2024★ 1 cited
TRRG: Towards Truthful Radiology Report Generation With Cross-modal Disease Clue Enhanced Large Language Model
Yuhao Wang, Chao Hao, Yawen Cui +4
The vision-language modeling capability of multi-modal large language models has attracted wide attention from the community. However, in medical domain, radiology report generatio…