15 papers
The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents
Zhexi Feng, Ruiyi Zhang, Yongbo Yang +1
A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can f…
Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators
Xinyue Zhao, Ruiyi Zhang, Liqin Ye +3
Nowcasting headline macroeconomic indicators, i.e., estimating an indicator's value for the current reference period before its official release, is critical for monetary policy an…
Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context
Zhexi Feng, Ruiyi Zhang, Yongbo Yang +1
Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expo…
ATLAS: Agentic Test-time Learning-to-Allocate Scaling
Peijia Qin, Qi Cao, Pengtao Xie
Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinemen…
AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models
Ruiyi Zhang, Peijia Qin, Qi Cao +2
AI models underpin data-centric applications from image and text processing to scientific discovery in biology, physics, and chemistry. Yet developing them remains heavily manual,…
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
Guiyao Tie, Jiawen Shi, Dingjie Song +20
Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, exper…