activity
20242026
most citedA Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond

2 citations · 2 across the 15 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

Jiachen Ma, Jiawen Zhang, Xiangtian Li +3

While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-…

cs.LG2026

M100: An Orchestrated Dataflow Architecture Powering General AI Computing

Yan Xie, Changkui Mao, Changsong Wu +34

As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based architectures offer versatility…

cs.LG2026

Native Reasoning Models: Training Language Models to Reason on Unverifiable Data

Yuanfu Wang, Zhixuan Liu, Xiangtian Li +2

The prevailing paradigm for training large reasoning models--combining Supervised Fine-Tuning (SFT) with Reinforcement Learning with Verifiable Rewards (RLVR)--is fundamentally con…

cs.LG2025

VRPRM: Process Reward Modeling via Visual Reasoning

Xinquan Chen, Chongying Yue, Bangwei Liu +3

Process Reward Model (PRM) is widely used in the post-training of Large Language Model (LLM) because it can perform fine-grained evaluation of the reasoning steps of generated cont…

cs.LG2025

Adversarial Preference Learning for Robust LLM Alignment

Yuanfu Wang, Pengyu Wang, Chenyang Xi +13

Modern language models often rely on Reinforcement Learning from Human Feedback (RLHF) to encourage safe behaviors. However, they remain vulnerable to adversarial attacks due to th…