activity
20212026
most citedHAGRID: A Human-LLM Collaborative Dataset for Generative Information-Seeking with Attribution

29 citations · 88 across the 13 of their papers we have counts for

collaborators

19 papers

cs.CR2026

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

Yongjie Wang, Xinyue Zhang, Kunhong Yao +4

Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such age…

cs.IR2026

LERA: LLM-Enhanced RAG for Ad Auction in Generative Chatbots

Haoran Sun, Xinrui Song, Xinyu Zhang +7

The integration of advertising auction mechanisms into large language model (LLM)-based chatbots presents a significant opportunity for commercialization, yet poses unique challeng…

cs.CL2026

Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven

Jiwei Tang, Zhijing Huang, Xinyu Zhang +5

Large Language Models (LLMs) have demonstrated exceptional performance across diverse tasks. However, their deployment in long-context scenarios faces high computational overhead a…

cs.CL2024

Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?

Crystina Zhang, Jing Lu, Vinh Q. Tran +3

Human understanding of text depends on general semantic concepts of words rather than their superficial forms. To what extent does our human intuition transfer to language models?…

cs.CV2024

Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation

Raphael Tang, Xinyu Zhang, Lixinyu Xu +5

Diffusion models are the state of the art in text-to-image generation, but their perceptual variability remains understudied. In this paper, we examine how prompts affect image var…

cs.CL2023

Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models

Xinyu Zhang, Sebastian Hofstätter, Patrick Lewis +2

Listwise rerankers based on large language models (LLM) are the zero-shot state-of-the-art. However, current works in this direction all depend on the GPT models, making it a singl…