23 papers
An AI Co-Data-Scientist for Prioritizing Candidate Biomarkers from Wearable Sensor Data
Yubin Kim, Salman Rahman, Samuel Schmidgall +33
Wearable devices generate continuous physiological and behavioral data, but converting these signals into clinically reviewable biomarker hypotheses remains labor-intensive. We int…
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
Palash Goyal, Mihir Parmar, Yiwen Song +3
The exponential growth of machine learning submissions has strained the traditional peer review process, resulting in slow feedback loops for authors and an immense burden on revie…
TeamBench: Evaluating Agent Coordination under Enforced Role Separation
Yubin Kim, Chanwoo Park, Taehan Kim +9
Agent systems often decompose a task across multiple roles, but these roles are typically specified by prompts rather than enforced by access controls. Without enforcement, a team…
TFRBench: A Reasoning Benchmark for Evaluating Forecasting Systems
Md Atik Ahamed, Mihir Parmar, Palash Goyal +7
We introduce TFRBench, the first benchmark designed to evaluate the reasoning capabilities of forecasting systems. Traditionally, time-series forecasting has been evaluated solely…
Watch and Learn: Learning to Use Computers from Online Videos
Chan Hee Song, Yiwen Song, Palash Goyal +4
Computer-using agents (CUAs) must plan task workflows across diverse and evolving applications, yet progress is limited by the lack of large-scale, high-quality training data. Exis…
HEART: Emotionally-Driven Test-Time Scaling of Language Models
Gabriela Pinto, Palash Goyal, Mihir Parmar +6
Test-time scaling has significantly improved how AI models solve problems, yet current methods often get stuck in repetitive, incorrect patterns of thought. We introduce HEART, a f…