4 papers
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
Enhao Huang, Pengyu Sun, Shuxun Wang +13
The Web3 ecosystem, underpinned by cryptographic primitives and decentralized consensus, represents a high-stakes environment where software vulnerabilities and incentive misalignm…
DecisionLLM: Large Language Models for Long Sequence Decision Exploration
Xiaowei Lv, Zhilin Zhang, Yijun Li +10
Long-sequence decision-making, which is usually addressed through reinforcement learning (RL), is a critical component for optimizing strategic operations in dynamic environments,…
Solving Formal Math Problems by Decomposition and Iterative Reflection
Yichi Zhou, Jianqiu Zhao, Yongxin Zhang +14
General-purpose Large Language Models (LLMs) have achieved remarkable success in intelligence, performing comparably to human experts on complex reasoning tasks such as coding and…
Multi-agent KTO: Reinforcing Strategic Interactions of Large Language Model in Language Game
Rong Ye, Yongxin Zhang, Yikai Zhang +3
Achieving Artificial General Intelligence (AGI) requires AI agents that can not only make stratigic decisions but also engage in flexible and meaningful communication. Inspired by…