collaborators

5 papers

cs.DB2026

The Table Says Otherwise: Testing LLMs with Counterfactual Relational Data

Xinzhi Wang, Chunwei Liu

Large language models (LLMs) are increasingly used to answer natural-language questions over structured data. However, when a table contains familiar real-world facts, it is unclea…

cs.DB2026

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing

Xinzhi Wang, Peter Baile Chen, Gerardo Vitagliano +5

Large language models with long context windows can answer complex questions directly from full-length academic, technical, and policy documents, but passing entire documents is of…

cs.CL2025

ACEBench: Who Wins the Match Point in Tool Usage?

Chen Chen, Xinlong Hao, Weiwen Liu +13

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex…

cs.CV2025

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

Zhihong Zhang, Xiaojian Huang, Jin Xu +4

Multimodal reward models (MRMs) play a crucial role in the training, inference, and evaluation of Large Vision Language Models (LVLMs) by assessing response quality. However, exist…

cs.LG2025

ToolACE: Winning the Points of LLM Function Calling

Weiwen Liu, Xu Huang, Xingshan Zeng +24

Function calling significantly extends the application boundary of large language models, where high-quality and diverse training data is critical for unlocking this capability. Ho…