3 papers
cs.CL2025
Efficient Long CoT Reasoning in Small Language Models
Zhaoyang Wang, Jinqi Jiang, Tian Qiu +3
Recent large reasoning models such as DeepSeek-R1 exhibit strong complex problems solving abilities by generating long chain-of-thought (CoT) reasoning steps. It is challenging to…
cs.LG2025
Anyprefer: An Agentic Framework for Preference Data Synthesis
Yiyang Zhou, Zhaoyang Wang, Tianle Wang +13
High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consum…
cs.AI2025
Synergistic Weak-Strong Collaboration by Aligning Preferences
Yizhu Jiao, Xuchao Zhang, Zhaoyang Wang +7
Current Large Language Models (LLMs) excel in general reasoning yet struggle with specialized tasks requiring proprietary or domain-specific knowledge. Fine-tuning large models for…