most citedThink Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking

2 citations · 3 across the 14 of their papers we have counts for

collaborators

16 papers

cs.CL2026

ASTRA: Automated Synthesis of agentic Trajectories and Reinforcement Arenas

Xiaoyu Tian, Haotian Wang, Shuaiting Chen +12

Large language models (LLMs) are increasingly used as tool-augmented agents for multi-step decision making, yet training robust tool-using agents remains challenging. Existing meth…

cs.AI2026

Coordinated Pandemic Control with Large Language Model Agents as Policymaking Assistants

Ziyi Shi, Xusen Guo, Hongliang Lu +7

Effective pandemic control requires timely and coordinated policymaking across administrative regions that are intrinsically interdependent. However, human-driven responses are oft…

stat.ML2026

Detecting Unobserved Confounders: A Kernelized Regression Approach

Yikai Chen, Yunxin Mao, Chunyuan Zheng +7

Detecting unobserved confounders is crucial for reliable causal inference in observational studies. Existing methods require either linearity assumptions or multiple heterogeneous…

cs.AI2025

When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs

Zhuoran Zhang, Tengyue Wang, Xilin Gong +4

Multimodal large language models (MLLMs) must resolve conflicts when different modalities provide contradictory information, a process we term modality following. Prior work measur…

cs.LG2025

Environment Inference for Learning Generalizable Dynamical System

Shixuan Liu, Yue He, Haotian Wang +4

Data-driven methods offer efficient and robust solutions for analyzing complex dynamical systems but rely on the assumption of I.I.D. data, driving the development of generalizatio…

cs.CV2025

BaseReward: A Strong Baseline for Multimodal Reward Model

Yi-Fan Zhang, Haihua Yang, Huanyu Zhang +11

The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for…