collaborators

8 papers

cs.LG2026

SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts

Bingshuai Liu, Ante Wang, Zijun Min +7

Large Language Models (LLMs) increasingly rely on reinforcement learning with verifiable rewards (RLVR) to elicit reliable chain-of-thought reasoning. However, the training process…

cs.LG2026

Orchestrating Tokens and Sequences: Dynamic Hybrid Policy Optimization for RLVR

Zijun Min, Bingshuai Liu, Ante Wang +4

Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising framework for optimizing large language models in reasoning tasks. However, existing RLVR algorithms focus…

cs.CL2026

Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation

Meiman Xiao, Ante Wang, Qingguo Hu +5

Precisely controlling the length of generated text is a common requirement in real-world applications. However, despite significant advancements in following human instructions, La…

cs.AI2025

BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

Yujie Lin, Jiayao Ma, Qingguo Hu +5

Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interventions often adopt a difference-…

cs.CL2025

LLM-OREF: An Open Relation Extraction Framework Based on Large Language Models

Hongyao Tu, Liang Zhang, Yujie Lin +4

The goal of open relation extraction (OpenRE) is to develop an RE model that can generalize to new relations not encountered during training. Existing studies primarily formulate O…

cs.CV2025

Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion

Qingguo Hu, Ante Wang, Jia Song +3

Large Vision-Language Models (LVLMs) have experienced significant advancements in recent years. However, their performance still falls short in tasks requiring deep visual percepti…