activity
20222025
most citedContrastive Learning Reduces Hallucination in Conversations

8 citations · 19 across the 17 of their papers we have counts for

collaborators

18 papers

cs.CL2025

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

Ziyou Hu, Zhengliang Shi, Minghang Zhu +5

Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both training and inference. However, exist…

cs.CL2025★ 1 cited

Bridging the Capability Gap: Joint Alignment Tuning for Harmonizing LLM-based Multi-Agent Systems

Minghang Zhu, Zhengliang Shi, Zhiwei Xu +5

The advancement of large language models (LLMs) has enabled the construction of multi-agent systems to solve complex tasks by dividing responsibilities among specialized agents, su…

cs.CL2025

Evolution without Large Models: Training Language Model with Task Principles

Minghang Zhu, Shen Gao, Zhengliang Shi +5

A common training approach for language models involves using a large-scale language model to expand a human-provided dataset, which is subsequently used for model training.This me…

cs.CV2025

Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization

Jiulong Wu, Zhengliang Shi, Shuaiqiang Wang +5

Large Visual Language Models (LVLMs) have demonstrated impressive capabilities across multiple tasks. However, their trustworthiness is often challenged by hallucinations, which ca…

cs.CL2025

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers

Zhengliang Shi, Lingyong Yan, Dawei Yin +3

Large language models (LLMs) have been widely integrated into information retrieval to advance traditional techniques. However, effectively enabling LLMs to seek accurate knowledge…

cs.CL2025

TTPA: Token-level Tool-use Preference Alignment Training Framework with Fine-grained Evaluation

Chengrui Huang, Shen Gao, Zhengliang Shi +2

Existing tool-learning methods usually rely on supervised fine-tuning, they often overlook fine-grained optimization of internal tool call details, leading to limitations in prefer…