activity
20242026
collaborators

20 papers

cs.CL2026

DocAtlas: Long-Document Understanding as Mutable-State Interaction

Hongchen Wei, Yuanzhe Wang, Bei Liu +8

Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually selec…

cs.CL2026

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding

Hongchen Wei, Yuanzhe Wang, Bei Liu +9

Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands o…

cs.CV2026

DiffCVE: Diffusion-based Compressed Video Enhancement

Wenqiang Xiao, Wenzhuo Ma, Junxi Zhang +1

Perceptual quality enhancement of severely compressed videos remains challenging due to complex artifact patterns and substantial information loss. Recent diffusion models have dem…

cs.CV2026

ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs

Yiling Gao, Hongchen Wei, Zhenzhong Chen

In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache overhead during inference. To a…

cs.CV2026

The First Challenge on Mobile Real-World Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

Jiatong Li, Zheng Chen, Kai Liu +91

This paper provides a review of the NTIRE 2026 challenge on mobile real-world image super-resolution, highlighting the proposed solutions and the resulting outcomes. The challenge…

cs.RO2026

Robotic Scene Cloning:Advancing Zero-Shot Robotic Scene Adaptation in Manipulation via Visual Prompt Editing

Binyuan Huang, Yuqing Wen, Yucheng Zhao +5

Modern robots can perform a wide range of simple tasks and adapt to diverse scenarios in the well-trained environment. However, deploying pre-trained robot models in real-world use…