6 citations · 6 across the 7 of their papers we have counts for
7 papers
HSRM: Hidden-State Reward Models for Test-Time Verification
Xianzhi Li, Xiaodan Zhu
Large language models can often generate plausible mathematical reasoning traces, but reliably identifying the correct solution among multiple candidates remains a key challenge. E…
Meta-Learning Reinforcement Learning for Crypto-Return Prediction
Junqiao Wang, Zhaoyang Guan, Guanyu Liu +7
Predicting cryptocurrency returns is notoriously difficult: price movements are driven by a fast-shifting blend of on-chain activity, news flow, and social sentiment, while labeled…
Detect, Explain, Escalate: Sustainable Dialogue Breakdown Management for LLM Agents
Abdellah Ghassel, Xianzhi Li, Xiaodan Zhu
Large Language Models (LLMs) have demonstrated substantial capabilities in conversational AI applications, yet their susceptibility to dialogue breakdowns poses significant challen…
Entropy-Gated Branching for Efficient Test-Time Reasoning
Xianzhi Li, Ethan Callanan, Abdellah Ghassel +1
Test-time compute methods can significantly improve the reasoning capabilities and problem-solving accuracy of large language models (LLMs). However, these approaches require subst…
Fine-Tuning Language Models with Differential Privacy through Adaptive Noise Allocation
Xianzhi Li, Ran Zmigrod, Zhiqiang Ma +2
Language models are capable of memorizing detailed patterns and information, leading to a double-edged effect: they achieve impressive modeling performance on downstream tasks with…
ChatGPT as Data Augmentation for Compositional Generalization: A Case Study in Open Intent Detection
Yihao Fang, Xianzhi Li, Stephen W. Thomas +1
Open intent detection, a crucial aspect of natural language understanding, involves the identification of previously unseen intents in user-generated text. Despite the progress mad…