11 papers
MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks
Yanjun Shao, Xiangru Tang, Jiwoong Sohn +10
Complex medical reasoning requires integrating heterogeneous clinical evidence across multiple inference steps. Large language models (LLMs) now approach this through two routes: i…
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering
Fangzhou Wu, Sandeep Silwal, Qiuyi Zhang
Query clustering organizes queries into groups that reflect shared latent capability demands, enabling capability-aware LLM evaluation. Existing clustering methods, which primarily…
DynMuon: A Dynamic Spectral Shaping View of Muon
Fangzhou Wu, Rikhav Shah, Sandeep Silwal +1
In recent years, Muon has emerged as the dominant method for training large language models, and transformers more broadly. The essential difference, when compared to standard grad…
SurfDesign: Effective Protein Design on Molecular Surfaces
Fang Wu, Shuting Jin, Xiangru Tang +5
Protein function is largely determined by molecular surface geometry and physicochemical complementarity, yet most protein design methods condition only on backbone structure. We i…
GeoCycler: Reward-Aligned 3D Diffusion for Constraint-Conditioned Cyclic Peptide Design
Jingjie Zhang, Hanqun Cao, Haosen Shi +10
Cyclic peptides are attractive therapeutic modalities because their closed-ring topology can improve stability and target specificity. However, de novo cyclic peptide design remain…
Latent Action Reparameterization for Efficient Agent Inference
Wenhao Huang, Qingwen Zeng, Qiyue Chen +11
Large language model (LLM) agents often rely on long sequences of low-level textual actions, resulting in large effective decision horizons and high inference cost. While prior wor…