Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
Bowen Liu, Shuning Wang, Xinpeng Ding +3
Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual…
cs.CV2026
MedHorizon: Towards Long-context Medical Video Understanding in the Wild
Bodong Du, Bowen Liu, Yang Yu +8
Medical multimodal large language models (MLLMs) have advanced image understanding and short-video analysis, but real clinical review often requires full-procedure video understand…
cs.CV2025
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
Honglong Yang, Shanshan Song, Yi Qin +6
Generalist Medical AI (GMAI) systems have demonstrated expert-level performance in biomedical perception tasks, yet their clinical utility remains limited by inadequate multi-modal…