3 papers
cs.AI2026
Toward Vibe Medicine: A Self-Evolving Multi-Agent Framework for Clinical Decision Support
Qianxue Zhang, Yiming Ren, Shihuan Qin +18
In recent years, the advances of large language models and autonomous agents have revolutionized the healthcare field, facilitating diagnosis and improving treatment results. Howev…
cs.CV2026
Multimodal Fusion for Sim2real Transfer in Visual Reinforcement Learning
Zichun Xu, Jingdong Zhao, Chenyu Guo +6
Depth information is robust to scene appearance variations and inherently carries 3D spatial details. Thus, a visual backbone based on the vision transformer is proposed to fuse RG…
cs.CV2025
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
Fufangchen Zhao, Liao Zhang, Daiqi Shi +5
We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to…