4 papers
Beyond Tier Labels: Role- and Deployment-Dependent Model Substitution in Multi-Call LLM Workflows
Renxiang Wang, Jiaming Cui
Large multi-call LLM systems pose a scientific problem that query-level routing does not capture: the value of a model depends on where it enters a dependent computation and on the…
Scaling Unverifiable Rewards: A Case Study on Visual Insights
Shuyu Gan, James Mooney, Pan Hao +4
Large Language Model (LLM) agents can increasingly automate complex reasoning through Test-Time Scaling (TTS), iterative refinement guided by reward signals. However, many real-wor…
A2P-Vis: an Analyzer-to-Presenter Agentic Pipeline for Visual Insights Generation and Reporting
Shuyu Gan, Renxiang Wang, James Mooney +1
Automating end-to-end data science pipeline with AI agents still stalls on two gaps: generating insightful, diverse visual evidence and assembling it into a coherent, professional…
Documentation Retrieval Improves Planning Language Generation
Renxiang Wang, Li Zhang
Certain strong LLMs have shown promise for zero-shot formal planning by generating planning languages like PDDL. Yet, the performance of most open-source models under 50B parameter…