collaborators

6 papers

cs.LG2026

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing

Guannan Lai, Haoran Hu, Long Chen +2

Existing LLM routing methods typically treat a model's single response to a query as its capability label for training routers. However, because LLM generation is inherently stocha…

cs.LG2026

Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark

Zhiqi Yu, Xingping Liu, Haobin Mao +4

Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, hand…

cs.CV2025

Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension

Lin Li, Wei Chen, Jiahui Li +2

Recent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relati…

cs.CL2025

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Jiahui Li, Lin Li, Tai-wei Chang +4

Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained…

cs.CV2025

Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards

Zijing Hu, Fengda Zhang, Long Chen +6

Diffusion models have achieved remarkable success in text-to-image generation. However, their practical applications are hindered by the misalignment between generated images and c…

cs.LG2025

Learning Causal Transition Matrix for Instance-dependent Label Noise

Jiahui Li, Tai-Wei Chang, Kun Kuang +3

Noisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of lea…