1 citations · 1 across the 6 of their papers we have counts for
8 papers
On the Paradoxical Interference between Instruction-Following and Task Solving
Yunjia Qi, Hao Peng, Xintong Shi +5
Instruction following aims to align Large Language Models (LLMs) with human intent by specifying explicit constraints on how tasks should be performed. However, we reveal a counter…
Evaluating Hydro-Science and Engineering Knowledge of Large Language Models
Shiruo Hu, Wenbo Shan, Yingjia Li +16
Hydro-Science and Engineering (Hydro-SE) is a critical and irreplaceable domain that secures human water supply, generates clean hydropower energy, and mitigates flood and drought…
WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
Guanzhong He, Zhen Yang, Jinxin Liu +3
Search agents have achieved significant advancements in enabling intelligent information retrieval and decision-making within interactive environments. Although reinforcement learn…
LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and Optimization
Xujia Wang, Yunjia Qi, Bin Xu
Parameter-Efficient Fine-Tuning (PEFT) methods, such as LoRA, significantly reduce the number of trainable parameters by introducing low-rank decomposition matrices. However, exist…
StoryWriter: A Multi-Agent Framework for Long Story Generation
Haotian Xia, Hao Peng, Yunjia Qi +4
Long story generation remains a challenge for existing large language models (LLMs), primarily due to two main factors: (1) discourse coherence, which requires plot consistency, lo…
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
Hao Peng, Yunjia Qi, Xiaozhi Wang +3
Reinforcement learning with verifiable rewards (RLVR) has become a key technique for enhancing large language models (LLMs), with verification engineering playing a central role. H…