1 citations · 1 across the 9 of their papers we have counts for
5 papers · 1 filter
Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation
Yirong Zeng, Zhang Sai, Yuxian Wang +4
Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fine-Tuning (SFT), which often lea…
TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist Ensembles
Yirong Zeng, Yufei Liu, Xiao Ding +9
Instruction Following (IF) is a core capability of LLMs, requiring strict adherence to diverse constraints, ranging from verifiable ones (e.g., output length) to unverifiable ones…
ToolACE-DEV: Self-Improving Tool Learning via Decomposition and EVolution
Xu Huang, Weiwen Liu, Xingshan Zeng +8
The tool-using capability of large language models (LLMs) enables them to access up-to-date external information and handle complex tasks. Current approaches to enhancing this capa…
iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use
Yirong Zeng, Xiao Ding, Yuxian Wang +8
Augmenting large language models (LLMs) with external tools is a promising approach to enhance their capabilities, especially for complex tasks. Synthesizing tool-use data through…
RU22Fact: Optimizing Evidence for Multilingual Explainable Fact-Checking on Russia-Ukraine Conflict
Yirong Zeng, Xiao Ding, Yi Zhao +5
Fact-checking is the task of verifying the factuality of a given claim by examining the available evidence. High-quality evidence plays a vital role in enhancing fact-checking syst…