4 papers
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
Wenhui Tan, Ruihua Song, Jiaze Li +2
Recent progress in multi-modal large language models (MLLMs) has significantly advanced video understanding. However, their performance on long-form videos remains limited by compu…
Xiaomi MiMo-VL-Miloco Technical Report
Jiaze Li, Jingyang Chen, Yuxun Qu +9
We open-source MiMo-VL-Miloco-7B and its quantized variant MiMo-VL-Miloco-7B-GGUF, a pair of home-centric vision-language models that achieve strong performance on both home-scenar…
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
Boshen Xu, Zihan Xiao, Jiaze Li +4
We introduce TimeViper, a hybrid vision-language model designed to tackle challenges of long video understanding. Processing long videos demands both an efficient model architectur…
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
Wenhui Tan, Jiaze Li, Jianzhong Ju +3
Large Language Models (LLMs) achieve superior performance through Chain-of-Thought (CoT) reasoning, but these token-level reasoning chains are computationally expensive and ineffic…