Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Rethinking Reward Models for Multi-Domain Test-Time Scaling
Dong Bok Lee, Seanie Lee, Sangwoo Park +12
The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning…
cs.AI2026
TS-Debate: Multimodal Collaborative Debate for Zero-Shot Time Series Reasoning
Patara Trirat, Jin Myung Kwak, Jay Heo +2
Recent progress at the intersection of large language models (LLMs) and time series (TS) analysis has revealed both promise and fragility. While LLMs can reason over temporal struc…