4 papers
HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning
Haowei Liu, Jiamian Wang, Hsin-Tai Wu +2
Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed. We introduce Hin…
RaCT: Ranking-aware Chain-of-Thought Optimization for LLMs
Haowei Liu, Xuyang Wu, Guohao Sun +2
In information retrieval, large language models (LLMs) have demonstrated remarkable potential in text reranking tasks by leveraging their sophisticated natural language understandi…
A Survey on Feedback-based Multi-step Reasoning for Large Language Models on Mathematics
Ting-Ruen Wei, Haowei Liu, Xuyang Wu +1
Recent progress in large language models (LLM) found chain-of-thought prompting strategies to improve the reasoning ability of LLMs by encouraging problem solving through multiple…
CLERF: Contrastive LEaRning for Full Range Head Pose Estimation
Ting-Ruen Wei, Haowei Liu, Huei-Chung Hu +3
We introduce a novel framework for representation learning in head pose estimation (HPE). Previously such a scheme was difficult due to head pose data sparsity, making triplet samp…