2 papers
cs.CV2026
Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models
Siyou Li, Huanan Wu, Juexi Shao +10
Despite the recent advances in the video understanding ability of multimodal large language models (MLLMs), long video understanding remains a challenge. One of the main issues is…
cs.LG2025
Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation
Siyou Li, Pengyao Qin, Huanan Wu +4
Automated radiology report generation (RRG) aims to produce detailed textual reports from clinical imaging, such as computed tomography (CT) scans, to improve the accuracy and effi…