activity
20242026
collaborators
Showing cs.CLShow all

15 papers · 1 filter

cs.CL2026

Data Selection for Multi-turn Dialogue Instruction Tuning

Bo Li, Shikun Zhang, Wei Ye

Instruction-tuned language models increasingly rely on large multi-turn dialogue corpora, but these datasets are often noisy and structurally inconsistent, with topic drift, repeti…

cs.CL2026

SteerRM: Debiasing Reward Models via Sparse Autoencoders

Mengyuan Sun, Zhuohao Yu, Weizheng Gu +2

Reward models (RMs) are critical components of alignment pipelines, yet they exhibit biases toward superficial stylistic cues, preferring better-presented responses over semantical…

cs.CL2026

Adaptive Test-Time Compute Allocation for Block Diffusion Language Models in Complex Reasoning

Yi Lu, Deyang Kong, Jianing Wang +8

Recent advances in block diffusion language models have demonstrated competitive performance and strong scalability on reasoning tasks. However, their test-time compute allocation…

cs.CL2026

ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback

Yutao Mou, Zhangchi Xue, Lijun Li +4

While LLM-based agents can interact with environments via invoking external tools, their expanded capabilities also amplify security risks. Monitoring step-level tool invocation be…

cs.CL2026

SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders

Zhuohao Yu, Xingru Jiang, Weizheng Gu +4

Watermarking LLM-generated text is critical for content attribution and misinformation prevention. However, existing methods compromise text quality, require white-box model access…

cs.CL2025

Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAG

Bo Li, Tian Tian, Zhenghua Xu +3

Dynamic retrieval-augmented generation (RAG) allows large language models (LLMs) to fetch external knowledge on demand, offering greater adaptability than static RAG. A central cha…