2 papers
cs.CV2026
EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports
Jianzhe Ma, Zhonghao Cao, Shangkui Chen +3
While video large language models (Video-LLMs) excel in understanding slow-paced, real-world egocentric videos, their capabilities in high-velocity, information-dense virtual envir…
cs.AI2026
Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support
Shuying Chen, Sen Cui, Zhong Cao
In this work, we propose Oph-Guid-RAG, a multimodal visual RAG system for ophthalmology clinical question answering and decision support. We treat each guideline page as an indepen…