5 papers
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors
Yifan Xu, Junren Chen, Yifan Chen
Reinforcement learning with verifiable rewards (RLVR) recently thrives in large language model (LLM) reasoning tasks. However, the reward sparsity and the long reasoning horizon ma…
Automated Quality Assessment of Blind Sweep Obstetric Ultrasound for Improved Diagnosis
Prasiddha Bhandari, Kanchan Poudel, Nishant Luitel +6
Blind Sweep Obstetric Ultrasound (BSOU) enables scalable fetal imaging in low-resource settings by allowing minimally trained operators to acquire standardized sweep videos for aut…
Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings
Yidong Jiang, Junrong Chen, Eftychia Makri +7
With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures. However, existing benc…
PAN: A World Model for General, Actionable, and Long-Horizon World Simulation
PAN Team, Zihan Liu, Yi Gu +12
A world model is a cognitive simulator of the real-world environment allowing biological agents to reason about how the world evolves, whether spontaneously or in response to their…
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
Qiyue Gao, Xinyu Pi, Kevin Liu +21
Internal world models (WMs) enable agents to understand the world's state and predict transitions, serving as the basis for advanced deliberative reasoning. Recent large Vision-Lan…