7 papers
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models
Faqiang Qian, Kang An, Weikun Zhang +6
Post-training alignment of large language models often combines supervised fine-tuning (SFT) on expert demonstrations with reinforcement learning (RL) from preference or verifiable…
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs
Ziliang Wang, Kang An, Faqiang Qian +5
Although reinforcement learning (RL) has expanded the cognitive boundaries of large language models (LLMs), it often remains vulnerable to the autoregressive curse in long-horizon…
Architecture Matters More Than Scale: A Comparative Study of Retrieval and Memory Augmentation for Financial QA Under SME Compute Constraints
Jianan Liu, Jing Yang, Xianyou Li +4
The rapid adoption of artificial intelligence (AI) and large language models (LLMs) is transforming financial analytics by enabling natural language interfaces for reporting, decis…
SELF-EMO: Emotional Self-Evolution from Recognition to Consistent Expression
Shaowei Zhang, Faqiang Qian, Yan Chen +5
Emotion Recognition in Conversation (ERC) has become a fundamental capability for large language models (LLMs) in human-centric interaction. Beyond accurate recognition, coherent e…
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
Ziliang Wang, Kang An, Xuhui Zheng +6
While search-augmented large language models (LLMs) exhibit impressive capabilities, their reliability in complex multi-hop reasoning remains limited. This limitation arises from t…
MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
Xuhui Zheng, Kang An, Ziliang Wang +3
Multimodal pre-training remains constrained by the descriptive bias of image-caption pairs, leading models to favor surface linguistic cues over grounded visual understanding. We i…