3 citations · 3 across the 4 of their papers we have counts for
6 papers
AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment
Zongqian Li, Yaoyiran Li, Yaohui Guo +3
Large language model agents can discover alphas, yet current methods have three weaknesses. The search cannot adapt during the run, automation usually ends at alpha generation whil…
WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance
Zhi Li, Tao Zhou, Yeqing Li +2
Delegating a web task involves more than asking a question; it requires transferring a policy: what to verify, how to handle uncertainty, which preferences matter, and when to stop…
ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders
Ofer Meshi, Krisztian Balog, Sally Goldman +5
The promise of LLM-based user simulators to improve conversational AI is hindered by a critical "realism gap," leading to systems that are optimized for simulated interactions, but…
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
Shenao Zhang, Yaqing Wang, Yinxiao Liu +5
Large Language Models (LLMs) trained via Reinforcement Learning (RL) have exhibited strong reasoning capabilities and emergent reflective behaviors, such as rethinking and error co…
Factored Agents: Decoupling In-Context Learning and Memorization for Robust Tool Use
Nicholas Roth, Christopher Hidey, Lucas Spangher +6
In this paper, we propose a novel factored agent architecture designed to overcome the limitations of traditional single-agent systems in agentic AI. Our approach decomposes the ag…
Project MPG: towards a generalized performance benchmark for LLM capabilities
Lucas Spangher, Tianle Li, William F. Arnold +6
There exists an extremely wide array of LLM benchmarking tasks, whereas oftentimes a single number is the most actionable for decision-making, especially by non-experts. No such ag…