3 papers
cs.CV2026
Online Self-Calibration Against Hallucination in Vision-Language Models
Minghui Chen, Chenxu Yang, Hengjie Zhu +3
Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image. Recent preference alignment…
cs.LG2026
Self-Distilled RLVR
Chenxu Yang, Chuanyu Qin, Qingyi Si +7
On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals…
cs.AI2025
Test-time Prompt Intervention
Chenxu Yang, Qingyi Si, Mz Dai +5
Test-time compute has led to remarkable success in the large language model (LLM) community, particularly for complex tasks, where longer chains of thought (CoTs) are generated to…