2 papers
cs.LG2026
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
Yi Jing, Zao Dai, Jinwu Hu +4
Model internals encode rich information about how a large language model (LLM) processes its training data; however, post-training data engineering largely relies on external signa…
cs.AI2026
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
Xiyuan Zhang, Huihang Wu, Jiayu Guo +8
We introduce FIRE, a comprehensive benchmark designed to evaluate both the theoretical financial knowledge of LLMs and their ability to handle practical business scenarios. For the…