papers

Publications (17)

cs.CV2023

NMS Threshold matters for Ego4D Moment Queries -- 2nd place solution to the Ego4D Moment Queries Challenge 2023

Lin Sui, Fangzhou Mu, Yin Li

This report describes our submission to the Ego4D Moment Queries Challenge 2023. Our submission extends ActionFormer, a latest method for temporal action localization. Our extensio…

cs.CV2026

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Zichao Lin, Yifeng Xie, Bowen Qu +30

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmar…

cs.CL2026

Kimi K2.5: Visual Agentic Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +339

We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that…

cs.CV2026

Affinity Contrastive Learning for Skeleton-based Human Activity Understanding

Hongda Liu, Yunfan Liu, Min Ren +3

In skeleton-based human activity understanding, existing methods often adopt the contrastive learning paradigm to construct a discriminative feature space. However, many of these a…

cs.CV2024

Harnessing Temporal Causality for Advanced Temporal Action Detection

Shuming Liu, Lin Sui, Chen-Lin Zhang +3

As a fundamental task in long-form video understanding, temporal action detection (TAD) aims to capture inherent temporal relations in untrimmed videos and identify candidate actio…

cs.CV2022

A Simple and Efficient Pipeline to Build an End-to-End Spatial-Temporal Action Detector

Lin Sui, Chen-Lin Zhang, Lixin Gu +1

Spatial-temporal action detection is a vital part of video understanding. Current spatial-temporal action detection methods mostly use an object detector to obtain person candidate…

cs.CV2026

StableMotion: One-Step Motion Estimation with Diffusion Prior

Ziyi Wang, Haipeng Li, Lin Sui +5

We present StableMotion, a novel framework that leverages geometric and content priors from pretrained large-scale image diffusion models for motion estimation in single-image rect…

cs.CV2026

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

Runjie Zhou, Youbo Shao, Haoyu Lu +16

We introduce WorldVQA, a benchmark designed to evaluate the atomic visual world knowledge of Multimodal Large Language Models (MLLMs). Unlike current evaluations, which often confl…

cs.CV2026

VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?

Yuanxin Liu, Kun Ouyang, Haoning Wu +7

Recent studies have shown that long chain-of-thought (CoT) reasoning can significantly enhance the performance of large language models (LLMs) on complex tasks. However, this benef…

cs.LG2026

Kimi K2: Open Agentic Intelligence

Kimi Team, Yifan Bai, Yiping Bao +195

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which…

cs.CV2025

Kimi-VL Technical Report

Kimi Team, Angang Du, Bohong Yin +92

We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong…

cs.CV2022

Salvage of Supervision in Weakly Supervised Object Detection

Lin Sui, Chen-Lin Zhang, Jianxin Wu

Weakly supervised object detection~(WSOD) has recently attracted much attention. However, the lack of bounding-box supervision makes its accuracy much lower than fully supervised o…

cs.CV2026

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24

VideoChat3 is a fully open, 4B-parameter video-centric multimodal large language model that combines an efficient Inflated 3D Vision Transformer and adaptive frame resolution with…

#video understanding#multimodal large language models#efficient video processing#data synthesis
cs.CL2026

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +398

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is…

cs.CV2026

Towards Pixel-Level VLM Perception via Simple Points Prediction

Tianhui Song, Haoyu Lu, Hao Yang +8

We present SimpleSeg, a strikingly simple yet highly effective approach to endow Multimodal Large Language Models (MLLMs) with native pixel-level perception. Our method reframes se…

cs.CV2025

TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos

Chen-Lin Zhang, Lin Sui, Shuming Liu +3

Temporal localization in untrimmed videos, which aims to identify specific timestamps, is crucial for video understanding but remains challenging. This task encompasses several sub…

cs.CL2026

Attention Residuals

Kimi Team, Guangyu Chen, Yu Zhang +34

Residual connections with PreNorm are standard in modern LLMs, yet they accumulate all layer outputs with fixed unit weights. This uniform aggregation causes uncontrolled hidden-st…