5 citations · 8 across the 14 of their papers we have counts for
28 papers · 1 filter
Unsupervised Post-Training of Foundation Models: A Survey
Yijie Xu, Qianyi Cai, Huizai Yao +9
Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearin…
Locally Confident, Globally Stuck: The Quality-Exploration Dilemma in Diffusion Language Models
Liancheng Fang, Aiwei Liu, Henry Peng Zou +7
Diffusion large language models (dLLMs) theoretically permit token decoding in arbitrary order, a flexibility that could enable richer exploration of reasoning paths than autoregre…
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
Leyi Pan, Shuchang Tao, Yunpeng Zhai +8
Reinforcement learning (RL) is pivotal for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, existing dLLM policy optimization methods suffe…
GenCNER: A Generative Framework for Continual Named Entity Recognition
Yawen Yang, Fukun Ma, Shiao Meng +2
Traditional named entity recognition (NER) aims to identify text mentions into pre-defined entity types. Continual Named Entity Recognition (CNER) is introduced since entity catego…
GapDNER: A Gap-Aware Grid Tagging Model for Discontinuous Named Entity Recognition
Yawen Yang, Fukun Ma, Shiao Meng +2
In biomedical fields, one named entity may consist of a series of non-adjacent tokens and overlap with other entities. Previous methods recognize discontinuous entities by connecti…
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
Yijie Xu, Huizai Yao, Zhiyu Guo +5
Large language models (LLMs) are increasingly deployed in specialized domains such as finance, medicine, and agriculture, where they face significant distribution shifts from their…