7 papers · 1 filter
Answer-Level Trust Selection for Physical Vision-Language Reasoning
Rongyu Yu, Ke Niu, Fengxiang He
Vision-language models (VLMs) can estimate physical quantities such as duration, speed, and acceleration from visual observations, but existing benchmarks primarily assess overall…
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
Kenton Tang, Yuzhu Chen, Fengxiang He
Alignment and adaptation in large language models heavily rely on reinforcement learning from human feedback (RLHF); yet, theoretical understanding of its generalisability remains…
PRISM: Parallel Reward Integration with Symmetry for MORL
Finn van der Knaap, Kejiang Qian, Zheng Xu +1
This work studies heterogeneous Multi-Objective Reinforcement Learning (MORL), where objectives can differ sharply in temporal frequency. Such heterogeneity allows dense objectives…
Rationality Measurement and Theory for Reinforcement Learning Agents
Kejiang Qian, Amos Storkey, Fengxiang He
This paper proposes a suite of rationality measures and associated theory for reinforcement learning agents, a property increasingly critical yet rarely explored. We define an acti…
DeXposure-FM: A Time-series, Graph Foundation Model for Credit Exposures and Stability on Decentralized Financial Networks
Aijie Shu, Wenbin Wu, Gbenga Ibikunle +1
Credit exposure in Decentralized Finance (DeFi) is often implicit and token-mediated, creating a dense web of inter-protocol dependencies. Thus, a shock to one token may result in…
DeXposure: A Dataset and Benchmarks for Inter-protocol Credit Exposure in Decentralized Financial Networks
Wenbin Wu, Kejiang Qian, Alexis Lui +5
We curate the DeXposure dataset, the first large-scale dataset for inter-protocol credit exposure in decentralized financial networks, covering global markets of 43.7 million entri…