activity
20242026
collaborators

9 papers

cs.CL2026

PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

Qiyao Wang, Xinyi Chen, Longze Chen +4

Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising application volumes. Prior benchmarks…

cs.AI2026

InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?

Qiyao Wang, Haoran Hu, Longze Chen +4

With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent-based project-level code sy…

cs.AI2026

FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration

Qiyao Wang, Hongbo Wang, Longze Chen +6

Scientific idea generation (SIG) is critical to AI-driven autonomous research, yet existing approaches are often constrained by a static retrieval-then-generation paradigm, leading…

cs.CL2026

Learning Ordinal Probabilistic Reward from Preferences

Longze Chen, Lu Wang, Renke Shan +6

Reward models are crucial for aligning large language models (LLMs) with human values and intentions. Existing approaches follow either Generative (GRMs) or Discriminative (DRMs) p…

cs.AI2026

Beyond Quantity: Trajectory Diversity Scaling for Code Agents

Guhong Chen, Chenghao Sun, Cheng Fu +16

As code large language models (LLMs) evolve into tool-interactive agents via the Model Context Protocol (MCP), their generalization is increasingly limited by low-quality synthetic…

cs.CL2025

IPBench: Benchmarking the Knowledge of Large Language Models in Intellectual Property

Qiyao Wang, Guhong Chen, Hongbo Wang +20

Intellectual Property (IP) is a highly specialized domain that integrates technical and legal knowledge, making it inherently complex and knowledge-intensive. Recent advancements i…