12 papers
REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
Zixing Chen, Xingyuan Liu, Jie Zhu +6
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and i…
Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue
Yi Wei, Shuo Jiang, Huaixia Dou +5
Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently multi-turn and path-dependent:…
FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions
Huaixia Dou, Jie Zhu, Minghao Wu +5
As large language models (LLMs) are increasingly deployed in financial services, a single non-compliant interaction can expose institutions to regulatory penalties and direct consu…
ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
Jie Zhu, Huaixia Dou, Shuo Jiang +5
Existing emotional support conversation (ESC) systems mainly rely on end-to-end response generation or coarse strategy supervision, offering limited interpretability and little sup…
Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
Jie Zhu, Yuanchen Zhou, Shuo Jiang +4
Process Reward Models (PRMs) supervise intermediate reasoning steps in large language models (LLMs), but existing PRMs are mainly trained on general-domain data and struggle with t…
Cross-Preference Learning for Sentence-Level and Context-Aware Machine Translation
Ying Li, Xinglin Lyu, Junhui Li +5
Context-aware machine translation (MT) leverages document-level information, yet it does not consistently outperform sentence-level MT, as contextual signals are unevenly beneficia…