2 papers
cs.LG2026
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
Muxi Diao, Lele Yang, Wuxuan Gong +6
Supervised Fine-Tuning (SFT) is the standard paradigm for domain adaptation, yet it frequently incurs the cost of catastrophic forgetting. In sharp contrast, on-policy Reinforcemen…
cs.CL2024
Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge Graph
Yutong Zhang, Lixing Chen, Shenghong Li +6
Large language models (LLMs) have demonstrated exceptional performance across a wide variety of domains. Nonetheless, generalist LLMs continue to fall short in reasoning tasks nece…