6 papers
Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation
Siyi Gu, Jialin Chen, Sophia Zhou +2
Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards. Distillation often relies on chain-of-t…
Reasoning through Verifiable Forecast Actions: Consistency-Grounded RL for Financial LLMs
Jialin Chen, Aosong Feng, Harshit Verma +7
Financial markets are characterized by extreme non-stationarity, low signal-to-noise ratios, and strong dependence on external information such as news, company fundamentals, and m…
LitBench: A Graph-Centric Large Language Model Benchmarking Tool For Literature Tasks
Andreas Varvarigos, Ali Maatouk, Jiasheng Zhang +4
While large language models (LLMs) have become the de facto framework for literature-related tasks, they still struggle to function as domain-specific literature agents due to thei…
Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
Yidong Jiang, Junrong Chen, Eftychia Makri +7
With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures. However, existing benc…
Multi-Modal Time Series Prediction via Mixture of Modulated Experts
Lige Zhang, Ali Maatouk, Jialin Chen +3
Real-world time series exhibit complex and evolving dynamics, making accurate forecasting extremely challenging. Recent multi-modal forecasting methods leverage textual information…
TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis
Austin Feng, Andreas Varvarigos, Ioannis Panitsas +7
Modern enterprises generate vast streams of time series metrics when monitoring complex systems, known as observability data. Unlike conventional time series from domains such as c…