1 citations · 1 across the 1 of their papers we have counts for
4 papers
Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
Yidong Jiang, Junrong Chen, Eftychia Makri +7
With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures. However, existing benc…
World Reasoning Arena
PAN Team, Qiyue Gao, Kun Zhou +15
World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM bench…
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
PAN Team, Jiannan Xiang, Yi Gu +31
A world model enables an intelligent agent to imagine, predict, and reason about how the world evolves in response to its actions, and accordingly to plan and strategize. While rec…
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
Qiyue Gao, Xinyu Pi, Kevin Liu +21
Internal world models (WMs) enable agents to understand the world's state and predict transitions, serving as the basis for advanced deliberative reasoning. Recent large Vision-Lan…