collaborators

6 papers

cs.CV2026

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs

Chengwen Liu, Hao Peng, Jisheng Dang +3

In multimodal video reasoning, reinforcement learning-based methods typically rely on simplistic and inflexible reasoning-length control strategies that fail to adapt to the model'…

cs.CL2026

Large Language Models Do Not Always Need Readable Language

Jiayi Zhu, Haoxuan Peng, Junxi Wang +3

Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is another model. This paper investigates whet…

cs.LG2026

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

Zhitong Wang, Songze Li, Hao Peng +4

Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents. However, conventional RL methods for long-horizon agentic tasks…

cs.CV2026

Powerful Teachers Matter: Text-Guided Multi-view Knowledge Distillation with Visual Prior Enhancement

Xin Zhang, Jianyang Xu, Hao Peng +5

Knowledge distillation transfers knowledge from large teacher models to smaller students for efficient inference. While existing methods primarily focus on distillation strategies,…

cs.CL2025

FactCheckmate: Preemptively Detecting and Mitigating Hallucinations in LMs

Deema Alnuhait, Neeraja Kirtane, Muhammad Khalifa +1

Language models (LMs) hallucinate. We inquire: Can we detect and mitigate hallucinations before they happen? This work answers this research question in the positive, by showing th…

cs.CL2025

LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language

Yubin Ge, Neeraja Kirtane, Hao Peng +1

As large language models (LLMs) have been deployed in various real-world settings, concerns about the harm they may propagate have grown. Various jailbreaking techniques have been…