activity
20242026
most citedOn The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability

6 citations · 7 across the 13 of their papers we have counts for

collaborators
Showing 2025Show all

6 papers · 1 filter

cs.CL2025

Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning

Qiang Liu, Wuganjing Song, Zhenzhou Lin +4

The reasoning capabilities of Large Language Models (LLMs) are typically developed through the single-turn reinforcement learning, whereas real-world applications often involve mul…

cs.CL2025

UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations

Qiuyang Lu, Fangjian Shen, Zhengkai Tang +4

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered fr…

cs.SE2025

SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development

Yaxin Du, Yuzhu Cai, Yifan Zhou +6

Large Language Models (LLMs) have shown strong capability in diverse software engineering tasks. However, feature-driven development, a highly prevalent real-world task that involv…

cs.AI2025★ 1 cited

Multi-agent Application System in Office Collaboration Scenarios

Songtao Sun, Jingyi Li, Yuanfei Dong +4

This paper introduces a multi-agent application system designed to enhance office collaboration efficiency and work quality. The system integrates artificial intelligence, machine…

cs.CL2025

AATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization

Junhui He, Junna Xing, Nan Wang +6

Long context large language models (LLMs) pose significant challenges for efficient serving due to the large memory footprint and high access overhead of KV cache. Retrieval-based…

cs.LG2025

PIPA: Preference Alignment as Prior-Informed Statistical Estimation

Junbo Li, Zhangyang Wang, Qiang Liu

Offline preference alignment for language models such as Direct Preference Optimization (DPO) is favored for its effectiveness and simplicity, eliminating the need for costly reinf…