2 citations · 3 across the 17 of their papers we have counts for
4 papers · 2 filters
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
Yu Liu, Zhilin Liu, Zhiwei Yang +7
As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception,…
MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
Yu Liu, Zhiwei Yang, Diandian Guo +7
Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information files (CIFs). Subtle chemical and structural…
Dialogue Model Optimization via Agent Game and Adaptive Tree-based GRPO
Kun Peng, Conghui Tan, Yu Liu +7
Open-ended dialogue agents aim to deliver engaging, personalized interactions by adapting to users' traits, but existing methods face critical limitations: over-reliance on pre-col…
PRISMA: Reinforcement Learning Guided Two-Stage Policy Optimization in Multi-Agent Architecture for Open-Domain Multi-Hop Question Answering
Yu Liu, Wenxiao Zhang, Cong Cao +10
Answering real-world open-domain multi-hop questions over massive corpora is a critical challenge in Retrieval-Augmented Generation (RAG) systems. Recent research employs reinforce…