activity
20182026
most citedMCare: Learning with Missing Modalities in Multimodal Healthcare Data

89 citations · 147 across the 25 of their papers we have counts for

collaborators
Showing cs.LGShow all

20 papers · 1 filter

cs.LG2026

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

Zhixin Zhang, Xinke Jiang, Zhibang Yang +5

Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is r…

cs.LG2026

The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment

Tianyu Jia, Yue Fang, Hongxin Ding +6

Process reward models (PRMs) enhance the reasoning capabilities of large language models (LLMs) by providing fine-grained feedback, yet training PRMs typically requires expensive s…

cs.LG2026

EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning

Hongxin Ding, Baixiang Huang, Yue Fang +6

Rubric-based rewards offer interpretable and fine-grained optimization signals for reinforcement learning in open-ended tasks where verifiable answers are unavailable. However, pre…

cs.LG2026

FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning

Yujie Feng, Hao Wang, Jian Li +6

Continual learning (CL) for large language models (LLMs) aims to enable sequential knowledge acquisition without catastrophic forgetting. Memory replay methods are widely used for…

cs.LG2025

Bridging Global Intent with Local Details: A Hierarchical Representation Approach for Semantic Validation in Text-to-SQL

Rihong Qiu, Zhibang Yang, Xinke Jiang +5

Text-to-SQL translates natural language questions into SQL statements grounded in a target database schema. Ensuring the reliability and executability of such systems requires vali…

cs.LG2025

ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning

Jinyang Zhang, Yue Fang, Hongxin Ding +5

Conventional continual pretraining (CPT) for large language model (LLM) domain adaptation often suffers from catastrophic forgetting and limited domain capacity. Existing strategie…