activity
20242026
collaborators

5 papers

cs.AI2026

FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

Ling Yue, Chaoqian Ouyang, Hang Xu +7

Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims difficult to verify. We present FactRevi…

cs.CV2026

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation

Longteng Guo, Xuanxu Lin, Dongze Hao +5

Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference across various subjects. Exis…

cs.CL2026

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

Bihui Yu, Caijun Jia, Jing Chi +6

Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, and multi-step reasoning. Cur…

cs.SE2025

Empowering RepoQA-Agent based on Reinforcement Learning Driven by Monte-carlo Tree Search

Guochang Li, Yuchen Liu, Zhen Qin +7

Repository-level software engineering tasks require large language models (LLMs) to efficiently navigate and extract information from complex codebases through multi-turn tool inte…

cs.SE2024

Lingma SWE-GPT: An Open Development-Process-Centric Language Model for Automated Software Improvement

Yingwei Ma, Rongyu Cao, Yongchang Cao +7

Recent advancements in LLM-based agents have led to significant progress in automatic software engineering, particularly in software maintenance and evolution. Despite these encour…