activity
20242026
collaborators

18 papers

cs.LG2026

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training

Binwen Tan, Jingchao Wang, Dengzhe Hou +6

Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits fin…

cs.AI2026

CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification

Dengzhe Hou, Lingyu Jiang, Fangzhou Lin +1

Electroencephalography (EEG) analysis in cognitive studies requires specialized expertise and involves many defensible choices over contrasts, channels, time windows, and statistic…

cs.CL2026

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

Dengzhe Hou, Lingyu Jiang, Fangzhou Lin +1

LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize…

cs.CV2026

CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

Fangzhou Lin, Peiran Li, Lingyu Xu +12

Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture th…

cs.AI2026

PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning

Lingyu Jiang, Zirui Li, Shuo Xing +6

The emergence of Large Reasoning Language Models (LRMs) has paved the way for tackling complex reasoning tasks through test-time scaling by generating long-form Chain-of-Thought (C…

cs.AI2026

CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning

Fangzhou Lin, Shuo Xing, Peiran Li +6

Parallel reasoning, where a generator samples many candidate solutions and an aggregator selects the best, is one of the most effective forms of test-time scaling in large language…