collaborators

7 papers

cs.IR2026

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

Bowen Qin, Yi Xie

Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the rep…

cs.CL2026

Hint Tuning: Less Data Makes Better Reasoners

Siqi Fan, Minghao Li, Xiaoqian Ma +6

Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning uniformly regardless of prob…

cs.SE2026

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis

Bowen Qin

CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to reduce the log to a manageable co…

cs.AI2026

BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions

Nan Huo, Xiaohan Xu, Jinyang Li +21

Large language models (LLMs) have demonstrated remarkable performance on single-turn text-to-SQL tasks, but real-world database applications predominantly require multi-turn intera…

cs.DB2026

SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications

Jinyang Li, Xiaolong Li, Ge Qu +17

Resolution of complex SQL issues persists as a significant bottleneck in real-world database applications. Current Large Language Models (LLMs), while adept at text-to-SQL translat…

cs.CL2025

Micro-Act: Mitigating Knowledge Conflict in LLM-based RAG via Actionable Self-Reasoning

Nan Huo, Jinyang Li, Bowen Qin +5

Retrieval-Augmented Generation (RAG) systems commonly suffer from Knowledge Conflicts, where retrieved external knowledge contradicts the inherent, parametric knowledge of large la…