3 papers
cs.CL2026
Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
Bingsen Chen, Boyan Li, Ping Nie +3
Existing benchmarks for Deep Research Agents (DRAs) treat report generation as a single-shot writing task, which fundamentally diverges from how human researchers iteratively draft…
cs.LG2025
Likert or Not: LLM Absolute Relevance Judgments on Fine-Grained Ordinal Scales
Charles Godfrey, Ping Nie, Natalia Ostapuk +3
Large language models (LLMs) obtain state of the art zero shot relevance ranking performance on a variety of information retrieval tasks. The two most common prompts to elicit LLM…
cs.LG2025
Nearest Neighbor Multivariate Time Series Forecasting
Huiliang Zhang, Ping Nie, Lijun Sun +1
Multivariate time series (MTS) forecasting has a wide range of applications in both industry and academia. Recently, spatial-temporal graph neural networks (STGNNs) have gained pop…