5 papers
Measuring Semantic Abstractness of SAE Features via Nonlocality
Chuqiao Lin, Shivaji Sondhi, Xiao-Liang Qi
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant a…
Agentic Publication Protocol: An Attempt to Modernize Scientific Publication
Sirui Lu, Xiao-Liang Qi
Scientific publication is still organized primarily around static manuscripts, even though much of scientific progress depends on tacit know-how: how to run code, reproduce figures…
QMBench: A Research Level Benchmark for Quantum Materials Research
Yanzhen Wang, Yiyang Jiang, Diana Golovanova +10
We introduce QMBench, a comprehensive benchmark designed to evaluate the capability of large language model agents in quantum materials research. This specialized benchmark assesse…
How Focused Are LLMs? A Quantitative Study via Repetitive Deterministic Prediction Tasks
Wanda Hou, Leon Zhou, Hong-Ye Hu +3
We investigate the performance of large language models on repetitive deterministic prediction tasks and study how the sequence accuracy rate scales with output length. Each such t…
AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems
Yang Li, Siqi Ping, Xiyu Chen +4
With the rapid progress of large language models (LLMs), LLM-powered multi-agent systems (MAS) are drawing increasing interest across academia and industry. However, many current M…