activity
20242026
collaborators

5 papers

cs.LG2026

The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment

Tianyu Jia, Yue Fang, Hongxin Ding +6

Process reward models (PRMs) enhance the reasoning capabilities of large language models (LLMs) by providing fine-grained feedback, yet training PRMs typically requires expensive s…

cs.LG2026

EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning

Hongxin Ding, Baixiang Huang, Yue Fang +6

Rubric-based rewards offer interpretable and fine-grained optimization signals for reinforcement learning in open-ended tasks where verifiable answers are unavailable. However, pre…

cs.CL2026

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning

Luan Zhang, Dandan Song, Zhijing Wu +8

Tool-integrated reasoning (TIR) enables large language models (LLMs) to enhance their capabilities by interacting with external tools, such as code interpreters (CI). Most recent s…

cs.CL2025

FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference

Runheng Liu, Xingchen Xiao, Heyan Huang +2

Retrieval-Augmented Language Modeling (RALM) by integrating large language models (LLM) with relevant documents from an external corpus is a proven method for enabling the LLM to g…

cs.CL2024

A Comprehensive Evaluation of Large Language Models on Aspect-Based Sentiment Analysis

Changzhi Zhou, Dandan Song, Yuhang Tian +6

Recently, Large Language Models (LLMs) have garnered increasing attention in the field of natural language processing, revolutionizing numerous downstream tasks with powerful reaso…