2 papers
cs.DB2026
MMTS-BENCH: A Comprehensive Benchmark for Time Series Understanding and Reasoning
Yao Yin, Zhenyu Xiao, Musheng Li +7
Time series data are central to domains such as finance, healthcare, and cloud computing, yet existing benchmarks for evaluating various large language models (LLMs) on temporal ta…
cs.CL2025
Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
Rheeya Uppaal, Apratim Dey, Yiting He +2
Recent alignment algorithms such as direct preference optimization (DPO) have been developed to improve the safety of large language models (LLMs) by training these models to match…