works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

RL Forgets! Towards Continual Policy Optimization

Mao-Lin Luo, Zhe-Xu Wang, Zi-Hao Zhou +4

The paper investigates catastrophic forgetting in continual post‑training of vision‑language models with reinforcement learning, introduces the MRCL benchmark, and proposes a repla…

cs.LG2026

Spectral Imbalance Causes Forgetting in Low-Rank Continual Adaptation

Hao Gu, Mao-Lin Luo, Zi-Hao Zhou +3

Parameter-efficient continual learning aims to adapt pre-trained models to sequential tasks without forgetting previously acquired knowledge. Most existing approaches treat continu…

cs.LG2026

Decouple then Converge: Handling Unknown Unlabeled Distributions in Long-Tailed Semi-Supervised Learning

Kai Gan, Tong Wei, Min-Ling Zhang

While long-tailed semi-supervised learning (LTSSL) has attracted growing attention in many real-world classification tasks, existing LTSSL algorithms typically assume that labeled…

cs.LG2026

DC-Merge: Improving Model Merging with Directional Consistency

Han-Chen Zhang, Zi-Hao Zhou, Mao-Lin Luo +3

Model merging aims to integrate multiple task-adapted models into a unified model that preserves the knowledge of each task. In this paper, we identify that the key to this knowled…

cs.LG2025

Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models

Jiajun Fan, Tong Wei, Chaoran Cheng +2

Balancing exploration and exploitation during reinforcement learning fine-tuning of generative models presents a critical challenge, as existing approaches rely on fixed divergence…

cs.LG2025

Tuning the Right Foundation Models is What you Need for Partial Label Learning

Kuang He, Wei Tang, Tong Wei +1

Partial label learning (PLL) seeks to train generalizable classifiers from datasets with inexact supervision, a common challenge in real-world applications. Existing studies have d…