7 papers
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
Zikang Liu, Junyi Li, Wayne Xin Zhao +3
Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) promise human-like interaction with software applications, yet long-horizon tasks remain c…
Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
Xuchen Pan, Yanxi Chen, Yushuo Chen +11
Trinity-RFT is a general-purpose, unified and easy-to-use framework designed for reinforcement fine-tuning (RFT) of large language models. It is built with a modular and decoupled…
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
Dawei Gao, Zitao Li, Yuexiang Xie +20
Driven by rapid advancements of Large Language Models (LLMs), agents are empowered to combine intrinsic knowledge with dynamic tool use, greatly enhancing their capacity to address…
GenSim: A General Social Simulation Platform with Large Language Model based Agents
Jiakai Tang, Heyang Gao, Xuchen Pan +11
With the rapid advancement of large language models (LLMs), recent years have witnessed many promising studies on leveraging LLM-based agents to simulate human social behavior. Whi…
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
Zikang Liu, Kun Zhou, Wayne Xin Zhao +3
Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language models (LVLMs). Despite the success,…
KIMAs: A Configurable Knowledge Integrated Multi-Agent System
Zitao Li, Fei Wei, Yuexiang Xie +6
Knowledge-intensive conversations supported by large language models (LLMs) have become one of the most popular and helpful applications that can assist people in different aspects…