5 papers
Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
Qingyu Ren, Qianyu He, Powei Chang +5
Language models often struggle to follow multi-constraint instructions that are crucial for real-world applications. Existing reinforcement learning (RL) approaches suffer from dep…
What Makes an Ideal Quote? Recommending "Unexpected yet Rational" Quotations via Novelty
Bowei Zhang, Jin Xiao, Guanglei Yue +4
Quotation recommendation aims to enrich writing by suggesting quotes that complement a given context, yet existing systems mostly optimize surface-level topical relevance and ignor…
Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following
Qingyu Ren, Qianyu He, Bowei Zhang +6
Reasoning models excel in complex problem solving but exhibit a concerning trade off between reasoning capabilities and instruction following abilities. Existing approaches for imp…
ChemHTS: Hierarchical Tool Stacking for Enhancing Chemical Agents
Zhucong Li, Jin Xiao, Bowei Zhang +5
Large Language Models (LLMs) have demonstrated remarkable potential in scientific research, particularly in chemistry-related tasks such as molecular design, reaction prediction, a…
QUILL: Quotation Generation Enhancement of Large Language Models
Jin Xiao, Bowei Zhang, Qianyu He +6
While Large language models (LLMs) have become excellent writing assistants, they still struggle with quotation generation. This is because they either hallucinate when providing f…