12 papers
Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
Yidong Jiang, Junrong Chen, Eftychia Makri +7
With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures. However, existing benc…
HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented Generation
Hiren Madhu, Ngoc Bui, Ali Maatouk +6
Embedding geometry plays a fundamental role in retrieval quality, yet dense retrievers for retrieval-augmented generation (RAG) remain largely confined to Euclidean space. However,…
TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis
Austin Feng, Andreas Varvarigos, Ioannis Panitsas +7
Modern enterprises generate vast streams of time series metrics when monitoring complex systems, known as observability data. Unlike conventional time series from domains such as c…
MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering
Jialin Chen, Aosong Feng, Ziyu Zhao +7
Understanding the relationship between textual news and time-series evolution is a critical yet under-explored challenge in applied data science. While multimodal learning has gain…
LitBench: A Graph-Centric Large Language Model Benchmarking Tool For Literature Tasks
Andreas Varvarigos, Ali Maatouk, Jiasheng Zhang +4
While large language models (LLMs) have become the de facto framework for literature-related tasks, they still struggle to function as domain-specific literature agents due to thei…
TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval
Jialin Chen, Ziyu Zhao, Gaukhar Nurbek +5
The ubiquity of dynamic data in domains such as weather, healthcare, and energy underscores a growing need for effective interpretation and retrieval of time-series data. These dat…