2 papers
cs.CV2026
Can Multimodal Large Language Models Truly Understand Small Objects?
Fujun Han, Junan Chen, Xintong Zhu +4
Multimodal Large Language Models (MLLMs) have shown promising potential in diverse understanding tasks, e.g., image and video analysis, math and physics olympiads. However, they re…
cs.CV2025
Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
Junan Chen, Trung Thanh Nguyen, Takahiro Komamizu +1
Recent advances in video captioning are driven by large-scale pretrained models, which follow the standard "pre-training followed by fine-tuning" paradigm, where the full model is…