collaborators

6 papers

cs.AI2026

Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning

Muyang Ye, Tian Lan, Feihu Jiang +10

Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains.…

cs.LG2026

EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning

Hongxin Ding, Baixiang Huang, Yue Fang +6

Rubric-based rewards offer interpretable and fine-grained optimization signals for reinforcement learning in open-ended tasks where verifiable answers are unavailable. However, pre…

cs.LG2026

GraphWalker: Patient Analogy Meets Information Gain for Clinical Reasoning with Large Language Models

Yue Fang, Weibin Liao, Yuxin Guo +8

Clinical reasoning over electronic health records (EHRs) is a fundamental yet challenging task in modern healthcare. While large language models (LLMs) offer a promising paradigm v…

cs.CL2026

The Tell-Tale Norm: Magnitude as a Signal for Reasoning Dynamics in Large Language Models

Jinyang Zhang, Hongxin Ding, Yue Fang +4

Recent work has sought to understand Large Language Models (LLMs) reasoning, yet a principled, model-intrinsic signal that captures its layer-wise reasoning dynamics remains undere…

cs.LG2025

ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning

Jinyang Zhang, Yue Fang, Hongxin Ding +5

Conventional continual pretraining (CPT) for large language model (LLM) domain adaptation often suffers from catastrophic forgetting and limited domain capacity. Existing strategie…

cs.LG2025

3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection

Hongxin Ding, Yue Fang, Runchuan Zhu +6

Large Language Models(LLMs) excel in general tasks but struggle in specialized domains like healthcare due to limited domain-specific knowledge.Supervised Fine-Tuning(SFT) data con…