activity
20242026
most citedUncovering Safety Risks of Large Language Models through Concept Activation Vector

2 citations · 2 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance

Jiachen Yu, Zhihao Xu, Junjie Wang +1

Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. Howeve…

cs.CL2026

Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text

Zhihao Xu, Rumei Li, Jiahuan Li +4

Enabling Large Language Models (LLMs) to effectively utilize tools in multi-turn interactions is essential for building capable autonomous agents. However, acquiring diverse and re…

cs.CL2025

Internal Value Alignment in Large Language Models through Controlled Value Vector Activation

Haoran Jin, Meng Li, Xiting Wang +4

Aligning Large Language Models (LLMs) with human values has attracted increasing attention since it provides clarity, transparency, and the ability to adapt to evolving scenarios.…

cs.CL2025

REWARD CONSISTENCY: Improving Multi-Objective Alignment from a Data-Centric Perspective

Zhihao Xu, Yongqi Tong, Xin Zhang +2

Multi-objective preference alignment in language models often encounters a challenging trade-off: optimizing for one human preference (e.g., helpfulness) frequently compromises oth…

cs.CL2024★ 2 cited

Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Zhihao Xu, Ruixuan Huang, Changyu Chen +1

Despite careful safety alignment, current large language models (LLMs) remain vulnerable to various attacks. To further unveil the safety risks of LLMs, we introduce a Safety Conce…