2 papers
cs.LG2026
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
Stephan Xie, Ben Cohen, Mononito Goswami +6
Time series question-answering (TSQA), in which we ask natural language questions to infer and reason about properties of time series, is a promising yet underexplored capability o…
cs.CL2024
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
Junhong Shen, Atishay Jain, Zedian Xiao +4
Large Language Model (LLM) agents are rapidly improving to handle increasingly complex web-based tasks. Most of these agents rely on general-purpose, proprietary models like GPT-4…