8 papers
Medical thinking with multiple images
Zonghai Yao, Benlu Wang, Yifan Zhang +8
Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a…
Rethinking Patient Education as Multi-turn Multi-modal Interaction
Zonghai Yao, Zhipeng Tang, Chengtao Lin +5
Most medical multimodal benchmarks focus on static tasks such as image question answering, report generation, and plain-language rewriting. Patient education is more demanding: sys…
Goal2Skill: Long-Horizon Manipulation with Adaptive Planning and Reflection
Zhen Liu, Xinyu Ning, Zhe Hu +8
Recent vision-language-action (VLA) systems have demonstrated strong capabilities in embodied manipulation. However, most existing VLA policies rely on limited observation windows…
Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models
Yuting Huang, Leilei Ding, Zhipeng Tang +7
While Vision-Language-Action (VLA) models hold promise in embodied intelligence, their large parameter counts lead to substantial inference latency that hinders real-time manipulat…
VLMPlanner: Integrating Visual Language Models with Motion Planning
Zhipeng Tang, Sha Zhang, Jiajun Deng +5
Integrating large language models (LLMs) into autonomous driving motion planning has recently emerged as a promising direction, offering enhanced interpretability, better controlla…
ElectricSight: 3D Hazard Monitoring for Power Lines Using Low-Cost Sensors
Xingchen Li, LiDian Wang, Yu Sheng +6
Protecting power transmission lines from potential hazards involves critical tasks, one of which is the accurate measurement of distances between power lines and potential threats,…