most citedExploiting Inherent Class Label: Towards Robust Scribble Supervised Semantic Segmentation

1 citations · 2 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CV2026

RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation

Shuang Zeng, Boxu Xie, Lei Zhu +6

Deep learning has greatly advanced medical image segmentation, but its success relies heavily on fully supervised learning, which requires dense annotations that are costly and tim…

cs.CV2026

Evaluation of Winning Solutions of 2025 Low Power Computer Vision Challenge

Zihao Ye, Yung-Hsiang Lu, Xiao Hu +14

The IEEE Low-Power Computer Vision Challenge (LPCVC) aims to promote the development of efficient vision models for edge devices, balancing accuracy with constraints such as latenc…

cs.CV2026

Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning

Zhengjian Yao, Yongzhi Li, Xinyuan Gao +3

We present "Narrative Weaver", a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual cont…

cs.CV2026

Bridging Degradation Discrimination and Generation for Universal Image Restoration

JiaKui Hu, Zhengjian Yao, Lujia Jin +1

Universal image restoration is a critical task in low-level vision, requiring the model to remove various degradations from low-quality images to produce clean images with rich det…

cs.CV2026

Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration

Zhengjian Yao, Jiakui Hu, Kaiwen Li +5

Blind face restoration remains a persistent challenge due to the inherent ill-posedness of reconstructing holistic structures from severely constrained observations. Current genera…

cs.CV2025

AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs

Xinliang Zhang, Lei Zhu, Hangzhou He +5

Multimodal Large Language Models (MLLMs) have demonstrated substantial value in unified text-image understanding and reasoning, primarily by converting images into sequences of pat…