From the 2 of 16 linked papers with an AI index.
16 papers
TAPO: Transition-Aware Policy Optimization for LLM Agents
Cong Li, Peixi Peng, Yisen Zhao +4
The paper introduces TAPO, a training framework that augments reinforcement learning for large language model agents with action‑conditioned next‑observation prediction, improving…
Kimi K3: Open Frontier Intelligence
Kimi Team, Tongtong Bai, Yifan Bai +398
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
Kai Chen, Zichen Ding, Jiaye Ge +22
The paper presents AgentCompass, an open‑source infrastructure that standardizes and simplifies the evaluation of large‑language‑model based autonomous agents by separating benchma…
Expectation Alignment of Language Models for Real-World User Expectations
Miaomiao Li, Yang Wang, Bin Liang +3
Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether they truly meet user expectations. Existing…
From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG
Guanhua Chen, Chuyue Huang, Yutong Yao +4
Multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities (entire images or scenes), creating a mismatch with fine-grained user queries and…
Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
Guanhua Chen, Yutong Yao, Shenghe Sun +5
Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA)…