2 papers
cs.AI2026
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
Cristina Garbacea, Heran Wang, Chenhao Tan
With the rise in capabilities of large language models (LLMs) and their deployment in real-world tasks, evaluating LLM alignment with human preferences has become an important chal…
cs.AI2025
StockMem: An Event-Reflection Memory Framework for Stock Forecasting
He Wang, Wenyilin Xiao, Songqiao Han +1
Stock price prediction is challenging due to market volatility and its sensitivity to real-time events. While large language models (LLMs) offer new avenues for text-based forecast…