collaborators

5 papers

cs.CV2026

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation

Qixiang Yin, Huanjin Yao, Yuchen Cai +5

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm by providing supervision on student-generated trajectories. However, existing OPD methods f…

cs.CL2025

ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning

Jianghao Chen, Wei Sun, Qixiang Yin +2

Long-form generation has become a critical and challenging application for Large Language Models (LLMs). Existing studies are limited by their reliance on scarce, high-quality long…

cs.AI2025

Towards Efficient Multimodal Unified Reasoning Model via Model Merging

Qixiang Yin, Huanjin Yao, Jianghao Chen +3

Although Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across diverse tasks, they encounter challenges in terms of reasoning efficiency, large…

cs.CL2025

LADM: Long-context Training Data Selection with Attention-based Dependency Measurement for LLMs

Jianghao Chen, Junhong Wu, Yangyifan Xu +1

Long-context modeling has drawn more and more attention in the area of Large Language Models (LLMs). Continual training with long-context data becomes the de-facto method to equip…

cs.CL2025

LR^2Bench: Evaluating Long-chain Reflective Reasoning Capabilities of Large Language Models via Constraint Satisfaction Problems

Jianghao Chen, Zhenlin Wei, Zhenjiang Ren +2

Recent progress in Large Reasoning Models (LRMs) has significantly enhanced the reasoning abilities of Large Language Models (LLMs), empowering them to tackle increasingly complex…