activity
20242026
collaborators
Showing cs.CLShow all

19 papers · 1 filter

cs.CL2026

EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection

Ancheng Xu, Zhihao Yang, Jingpeng Li +9

E-commerce platforms increasingly rely on Large Language Models (LLMs) and Vision Language Models (VLMs) to detect illicit or misleading product content. However, these models rema…

cs.CL2026

PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

Qiyao Wang, Xinyi Chen, Longze Chen +4

Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising application volumes. Prior benchmarks…

cs.CL2026

DEFT: Distribution-guided Efficient Fine-Tuning for Human Alignment

Liang Zhu, Feiteng Fang, Yuelin Bai +4

Reinforcement Learning from Human Feedback (RLHF), using algorithms like Proximal Policy Optimization (PPO), aligns Large Language Models (LLMs) with human values but is costly and…

cs.CL2026

RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning

Yukun Chen, Jiaming Li, Longze Chen +10

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a prevailing paradigm for enhancing reasoning in Multimodal Large Language Models (MLLMs). However, relying sol…

cs.CL2026

Learning Ordinal Probabilistic Reward from Preferences

Longze Chen, Lu Wang, Renke Shan +6

Reward models are crucial for aligning large language models (LLMs) with human values and intentions. Existing approaches follow either Generative (GRMs) or Discriminative (DRMs) p…

cs.CL2026

Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR

Jiaming Li, Longze Chen, Ze Gong +5

Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have empowered large language models (LLMs) to tackle challenging reasoning tasks such as mathematics and p…