2 papers
cs.LG2026
Scaling Reward Modeling without Human Supervision
Jingxuan Fan, Yueying Li, Zhenting Qi +4
Learning from feedback is an instrumental process for advancing the capabilities and safety of frontier models, yet its effectiveness is often constrained by cost and scalability.…
q-fin.TR2025
MountainLion: A Multi-Modal LLM-Based Agent System for Interpretable and Adaptive Financial Trading
Siyi Wu, Junqiao Wang, Zhaoyang Guan +11
Cryptocurrency trading is a challenging task requiring the integration of heterogeneous data from multiple modalities. Traditional deep learning and reinforcement learning approach…