works on

From the 1 of 12 linked papers with an AI index.

most citedVLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

3 citations · 3 across the 4 of their papers we have counts for

collaborators

12 papers

cs.CL2026

Training Skills Like Parameters via Self-Supervised Semantic Diffusion

Mo Li, Zixin Yin, Ting Cao +1

The paper introduces a self‑supervised framework that lets a language model acquire and store textual skills in an external library using diffusion‑style reconstruction loss, witho…

cs.LG2026

GatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series Forecasting

Qitai Tan, Ruiwen Gu, Yilin Su +3

Time series forecasting requires models to capture diverse, often mutually exclusive, temporal dynamics, from smooth trend continuation to nonstationary drift and strict phase-alig…

cs.SE2026

Learning from Execution: Self-Evolving Memory for Private-Library Code Generation

Mofei Li, Taozhi Chen, Guowei Yang +1

Large Language Models (LLMs) have achieved strong performance on general code generation, but their effectiveness drops sharply in enterprise settings where software development re…

cs.CV20263 cited

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Haodong Duan, Xinyu Fang, Junming Yang +41

We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friendly and comprehensive framework f…

cs.AI2026

ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks

Qitai Tan, Zefang Zong, Yang Li +3

Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teac…

cs.CV2026

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Sihan Yang, Runsen Xu, Yiman Xie +10

Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relati…