3 papers
cs.AI2026
When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets
Zeshi Dai, Zimo Peng, Zerui Cheng +1
We present CAIA, a benchmark exposing a critical blind spot in AI evaluation: the inability of state-of-the-art models to operate in adversarial, high-stakes environments where mis…
cs.CL2025
Semantic Self-Consistency: Enhancing Language Model Reasoning via Semantic Weighting
Tim Knappe, Ryan Li, Ayush Chauhan +3
While large language models (LLMs) have rapidly improved their performance on a broad number of tasks, they still often fall short on reasoning tasks. As LLMs become more integrate…
cs.LG2024
FedGAT: A Privacy-Preserving Federated Approximation Algorithm for Graph Attention Networks
Siddharth Ambekar, Yuhang Yao, Ryan Li +1
Federated training methods have gained popularity for graph learning with applications including friendship graphs of social media sites and customer-merchant interaction graphs of…