1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.AI2025
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
Hao Chen, Zhexin Hu, Jiajun Chai +7
Training LLMs to invoke tools and leverage retrieved information necessitates high-quality, diverse data. However, existing pipelines for synthetic data generation often rely on te…
cs.AI2025★ 1 cited
Adaptive Selection of Symbolic Languages for Improving LLM Logical Reasoning
Xiangyu Wang, Haocheng Yang, Fengxiang Cheng +1
Large Language Models (LLMs) still struggle with complex logical reasoning. While previous works achieve remarkable improvements, their performance is highly dependent on the corre…
cs.LG2025
Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling
Han Qi, Haochen Yang, Qiaosheng Zhang +1
We study the problem of reinforcement learning from human feedback (RLHF), a critical problem in training large language models, from a theoretical perspective. Our main contributi…