5 papers
DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation
Xifeng Xue, Xiaokang Wang, Zirui Li +2
Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as…
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation
Yue Xun, Junyu Liu, Qian Niu +10
We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese M…
Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure
Zirui Li, Xuefeng Bai, Kehai Chen +4
Latent or continuous chain-of-thought methods replace explicit textual rationales with a number of internal latent steps, but these intermediate computations are difficult to evalu…
PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning
Lingyu Jiang, Zirui Li, Shuo Xing +6
The emergence of Large Reasoning Language Models (LRMs) has paved the way for tackling complex reasoning tasks through test-time scaling by generating long-form Chain-of-Thought (C…
DocMMIR: A Framework for Document Multi-modal Information Retrieval
Zirui Li, Siwei Wu, Yizhi Li +3
The rapid advancement of unsupervised representation learning and large-scale pre-trained vision-language models has significantly improved cross-modal retrieval tasks. However, ex…