18 papers
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training
Binwen Tan, Jingchao Wang, Dengzhe Hou +6
Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits fin…
CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification
Dengzhe Hou, Lingyu Jiang, Fangzhou Lin +1
Electroencephalography (EEG) analysis in cognitive studies requires specialized expertise and involves many defensible choices over contrasts, channels, time windows, and statistic…
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
Dengzhe Hou, Lingyu Jiang, Fangzhou Lin +1
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize…
CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences
Fangzhou Lin, Peiran Li, Lingyu Xu +12
Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edits and do not fully capture th…
PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning
Lingyu Jiang, Zirui Li, Shuo Xing +6
The emergence of Large Reasoning Language Models (LRMs) has paved the way for tackling complex reasoning tasks through test-time scaling by generating long-form Chain-of-Thought (C…
CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning
Fangzhou Lin, Shuo Xing, Peiran Li +6
Parallel reasoning, where a generator samples many candidate solutions and an aggregator selects the best, is one of the most effective forms of test-time scaling in large language…