14 papers
Persistent Object Narratives for Token-Efficient Video Language Models
Junzhe Chen, Siyuan Meng, Xiaojie Guo
Video large language models (Video-LLMs) have made strong progress in open-ended video understanding. However, their visual interfaces remain token-intensive and provide limited ex…
Fabric Image Demoiréing Benchmark from Synthesis to Restoration
Pengchao Wei, Xiaojie Guo
Fabric moiré is a sampling-induced aliasing artifact caused by the interaction between fine textile patterns and camera sensor grids, producing structured interference that severe…
Principled Reflection Separation via Nonlinear Superposition and Feature Interaction
Qiming Hu, Mingjia Li, Yuntong Li +1
Single-image reflection separation is fundamentally challenged by the entanglement of transmission and reflection layers under complex image formation processes. Existing approache…
Representative Attention For Vision Transformers
Yuntong Li, Hainuo Wang, Hengxing Liu +2
Linear attention has emerged as a promising direction for scaling Vision Transformers beyond the quadratic cost of dense self-attention. A prevalent strategy is to compress spatial…
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
Junzhe Chen, Siyuan Meng, Yuxi Chen +3
Video large language models (Video-LLMs) have made strong progress in general video understanding, but their ability to maintain temporal object consistency remains underexplored.…
On the Global Photometric Alignment for Low-Level Vision
Mingjia Li, Tianle Du, Hainuo Wang +2
Supervised low-level vision models rely on pixel-wise losses against paired references, yet paired training sets exhibit per-pair photometric inconsistency, say, different image pa…