activity
20222025
most citedGenerative Entity Typing with Curriculum Learning

1 citations · 1 across the 5 of their papers we have counts for

collaborators

11 papers

cs.AI2025

ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints

Rui Xu, Dakuan Lu, Zicheng Zhao +5

Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. Ho…

cs.CL2025

Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges

Weiyuan Li, Xintao Wang, Siyu Yuan +5

As large language models (LLMs) grow more capable, they face increasingly diverse and complex tasks, making reliable evaluation challenging. The paradigm of LLMs as judges has emer…

cs.CL2025

Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games

Yikai Zhang, Ye Rong, Siyu Yuan +3

Existing language agents often encounter difficulties in dynamic adversarial games due to poor strategic reasoning. To mitigate this limitation, a promising approach is to allow ag…

cs.CL2025

ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models

Qianyu He, Siyu Yuan, Xuefeng Li +2

Large language models (LLMs) with chain-of-thought reasoning have demonstrated remarkable problem-solving capabilities, but controlling their computational effort remains a signifi…

cs.CL2025

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles

Jiangjie Chen, Qianyu He, Siyu Yuan +9

Large Language Models (LLMs), such as OpenAI's o1 and DeepSeek's R1, excel at advanced reasoning tasks like math and coding via Reinforcement Learning with Verifiable Rewards (RLVR…

cs.CL2025

ARIA: Training Language Agents with Intention-Driven Reward Aggregation

Ruihan Yang, Yikai Zhang, Aili Chen +5

Large language models (LLMs) have enabled agents to perform complex reasoning and decision-making through free-form language interactions. However, in open-ended language action en…