5 papers
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training
Binwen Tan, Jingchao Wang, Dengzhe Hou +6
Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits fin…
CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification
Dengzhe Hou, Lingyu Jiang, Fangzhou Lin +1
Electroencephalography (EEG) analysis in cognitive studies requires specialized expertise and involves many defensible choices over contrasts, channels, time windows, and statistic…
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
Dengzhe Hou, Lingyu Jiang, Fangzhou Lin +1
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize…
Physics-Aware Video Instance Removal Benchmark
Zirui Li, Xinghao Chen, Lingyu Jiang +5
Video Instance Removal (VIR) requires removing target objects while maintaining background integrity and physical consistency, such as specular reflections and illumination interac…
WMF-AM: Probing LLM Working Memory via Depth-Parameterized Cumulative State Tracking
Dengzhe Hou, Lingyu Jiang, Deng Li +3
Existing large language models (LLMs) evaluations use fixed-difficulty benchmarks that cannot adapt as models improve, and rarely isolate specific cognitive processes. We introduce…