activity
20242026
most citedAdvancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning

6 citations · 10 across the 44 of their papers we have counts for

collaborators

53 papers

cs.CL2026

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

Dingzirui Wang, Xuanliang Zhang, Keyan Xu +2

Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research sett…

cs.CL2026

SDAM: Structure-Difference-Aware Memory Evolution for Complex Text-to-SQL

Keyan Xu, Dingzirui Wang, Xuanliang Zhang +2

Text-to-SQL aims to convert natural language questions into executable SQL queries. While memory-based agent system improves complex SQL generation, existing memory design neglect…

cs.CL2026

CurateEvo: Data-Curation Evolving for Agentic Post-Training

Dingzirui Wang, Xuanliang Zhang, Keyan Xu +2

Large language model (LLM) agents require post-training methods that can improve long-horizon decision making from environment feedback. However, existing agentic post-training pip…

cs.SE2026

Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards

Tianhao Niu, Ziyu Han, Qiguang Chen +5

Automatic data visualization generation has advanced rapidly with multi-modal large language models, yet existing efforts largely focus on static charts and overlook the interactiv…

cs.CL2026

What You See Is What You Get: Observation-Aligned Supervision for Chart-to-Code Generation

Tianhao Niu, Qingfu Zhu, Wanxiang Che

Chart-to-code generation is commonly trained through supervised fine-tuning on reference plotting scripts, implicitly treating the gold code as a fully observable target. However,…

cs.CL2026

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape

Hao Chen, Ziyu Han, Yukun Yan +3

As Large Language Models (LLMs) advance toward open-ended autonomous agents, the mechanisms used to evaluate and guide their behavior must evolve accordingly. This work introduces…