43 papers
SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks
Jinchao Hu, Meizhi Zhong, Kehai Chen +1
Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries. This is especially important in open-domain…
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
Bingzheng Qu, Kehai Chen, Xuefeng Bai +1
Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-sp…
Beyond Rigid: Benchmarking Non-Rigid Video Editing
Bingzheng Qu, Xuefeng Bai, Kehai Chen +1
As video generation models are increasingly expected to manipulate physical dynamics, there is a growing need to move evaluation beyond appearance fidelity and semantic alignment.…
Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure
Zirui Li, Xuefeng Bai, Kehai Chen +4
Latent or continuous chain-of-thought methods replace explicit textual rationales with a number of internal latent steps, but these intermediate computations are difficult to evalu…
Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware Guidance
Yingjie Zhu, Xuefeng Bai, Kehai Chen +4
Reasoning over table images remains challenging for Large Vision-Language Models (LVLMs) due to complex layouts and tightly coupled structure-content information. Existing solution…
Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning
Hongbin Zhang, Chaozheng Wang, Kehai Chen +4
On-policy self-distillation (OPSD) is an emerging LLM post-training paradigm in which the model serves as its own teacher: conditioned on privileged information such as a reference…