activity
20242026
most citedScaling Spike-driven Transformer with Efficient Spike Firing Approximation Training

48 citations · 49 across the 11 of their papers we have counts for

collaborators

17 papers

cs.CV2026

RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation

Shuang Zeng, Boxu Xie, Lei Zhu +6

Deep learning has greatly advanced medical image segmentation, but its success relies heavily on fully supervised learning, which requires dense annotations that are costly and tim…

cs.CV2026

HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization

Xuerui Qiu, Yutao Cui, Guozhen Zhang +9

Unified Multimodal Models struggle to bridge the fundamental gap between the abstract representations needed for visual understanding and the detailed primitives required for gener…

cs.CV2026

Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context

JiaKui Hu, Jialun Liu, Liying Yang +7

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory f…

cs.CV2026

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

Jialun Liu, Tian Li, Xiao Cao +20

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and r…

cs.CV2026

Spatial-Temporal State Propagation Autoregressive Model for 4D Object Generation

Liying Yang, Jialun Liu, Jiakui Hu +5

Generating high-quality 4D objects with spatial-temporal consistency is still formidable. Existing diffusion-based methods often struggle with spatial-temporal inconsistency, as th…

cs.CV2026

Bridging Degradation Discrimination and Generation for Universal Image Restoration

JiaKui Hu, Zhengjian Yao, Lujia Jin +1

Universal image restoration is a critical task in low-level vision, requiring the model to remove various degradations from low-quality images to produce clean images with rich det…