9 papers
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
Shuhan Xue, Zixin Ding, Yichen Shen +6
Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability beca…
Q-Score: A Quantum-Native Scoring Function for Molecular Docking
Kangyu Zheng, Yidong Zhou, Ruihao Li +3
Molecular docking predicts how a small molecule binds to a protein and is a key bottleneck in drug discovery. Classical scoring functions sum empirical pairwise contacts, blind to…
Scaling Textual Gradients via Sampling-Based Momentum
Zixin Ding, Junyuan Hong, Zhan Shi +6
LLM-based prompt optimization, which uses LLM-provided ``textual gradients'' (feedback) to refine prompts, has emerged as an effective method for automatic prompt engineering. Howe…
Learning to Trigger: Reinforcement Learning at the Large Hadron Collider
Zixin Ding, Shaghayegh Emami, Giovanna Salvi +7
High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tight constraints on bandwidth, latency, and…
AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
Bo Zhang, Tzu-Yen Ma, Zichen Tang +18
We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS features three key advances: (1) D…
For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
Wenlong Deng, Qi Zeng, Jiaming Zhang +5
Data valuation is essential for enhancing the transparency and accountability of large language models (LLMs) and vision-language models (VLMs). However, existing methods typically…