3 papers
cs.CV2026
Variance & Greediness: A comparative study of metric-learning losses
Donghuo Zeng, Hao Niu, Zhi Li +1
Metric learning is central to retrieval, yet its effects on embedding geometry and optimization dynamics are not well understood. We introduce a diagnostic framework, VARIANCE (int…
cs.IR2025
An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
Zhi Li, Yanan Wang, Hao Niu +2
Multimodal large language models have recently achieved remarkable progress in video question answering (VideoQA) by jointly processing visual, textual, and audio information. Howe…
cs.CV2025
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
Yanan Wang, Julio Vizcarra, Zhi Li +2
Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fi…