activity
20242026
most citedInference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2026

Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization

Qingchan Zhu, Weihang You, Hanqi Jiang +3

Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-s…

cs.CV2026

ViThinker: Active Vision-Language Reasoning via Dynamic Perceptual Querying

Weihang You, Qingchan Zhu, David Liu +3

Chain-of-Thought (CoT) reasoning excels in language models but struggles in vision-language models due to premature visual-to-text conversion that discards continuous information s…

cs.CL2025

AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models

Qin Zhu, Fei Huang, Runyu Peng +6

While logical reasoning evaluation of Large Language Models (LLMs) has attracted significant attention, existing benchmarks predominantly rely on multiple-choice formats that are v…

cs.CL20241 cited

Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation

Qin Zhu, Qingyuan Cheng, Runyu Peng +5

The training process of large language models (LLMs) often involves varying degrees of test data contamination. Although current LLMs are achieving increasingly better performance…

cs.CL2024

Unified Active Retrieval for Retrieval Augmented Generation

Qinyuan Cheng, Xiaonan Li, Shimin Li +7

In Retrieval-Augmented Generation (RAG), retrieval is not always helpful and applying it to every instruction is sub-optimal. Therefore, determining whether to retrieve is crucial…