Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
Chao Chen, Zhixin Ma, Yongqi Li +4
Multimodal reasoning aims to enhance the capabilities of MLLMs by incorporating intermediate reasoning steps before reaching the final answer. It has evolved from text-only reasoni…
cs.CV2025
RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
Meilong Xu, Di Fu, Jiaxing Zhang +7
Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particul…