activity
20242026
collaborators

14 papers

cs.AI2026

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

Tianyu Huai, Tingshuo Fan, Xinchi Chen +5

As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmark…

cs.LG2026

ABOPD: Antibody CDR Design via On-Policy Distillation

Zhuo Yang, Jiaying He, Jiaqing Xie +5

Antibodies are essential therapeutic molecules, and their complementarity-determining regions (CDRs) form the primary antigen-recognition interface. Recent protein generative model…

cs.LG2026

Eigenbasis-Independent Learnable Spectral Positional Encodings for Directed Graphs via Hermitian Block Krylov Subspaces

Jiaqing Xie, Yuxin Wang

Spectral positional encodings (PEs) for \emph{directed} graphs face two obstacles: magnetic Laplacians require an Hermitian eigendecomposition per potential, and their com…

cs.CL2026

Rethinking Scientific Discovery in the Agentic Era

Yining Zheng, Yuxin Wang, Jiahao Lu +27

Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem formulation, literature gro…

cs.CL2026

AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering

Yuxin Wang, Jiahao Lu, Qifeng Wu +5

Large Language Models (LLMs) have achieved remarkable performance in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, this approach often leads to ``over-…

cs.CL2026

AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts

Shicheng Fang, Yuxin Wang, Xiaoran Liu +6

The evolution of Large Language Models (LLMs) into autonomous agents necessitates the management of extensive, dynamic contexts. Current benchmarks, however, remain largely static,…