4 citations · 7 across the 13 of their papers we have counts for
15 papers · 1 filter
Test-Time Scaling for Video Diffusion Models via Diagnosis-Guided Candidate Recycling
Hangzhou He, Lunhao Duan, Shanshan Zhao +4
Recent video diffusion models have achieved remarkable generation quality, but high-fidelity results still largely depend on closed-source systems or costly large-scale infrastruct…
Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration
Zhengjian Yao, Jiakui Hu, Kaiwen Li +5
Blind face restoration remains a persistent challenge due to the inherent ill-posedness of reconstructing holistic structures from severely constrained observations. Current genera…
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
Xinliang Zhang, Lei Zhu, Hangzhou He +5
Multimodal Large Language Models (MLLMs) have demonstrated substantial value in unified text-image understanding and reasoning, primarily by converting images into sequences of pat…
Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
Hangzhou He, Lei Zhu, Kaiwen Li +5
Concept Bottleneck Models (CBMs) provide inherent interpretability by first predicting a set of human-understandable concepts and then mapping them to labels through a simple class…
Inter- and Intra-image Refinement for Few Shot Segmentation
Ourui Fu, Hangzhou He, Kaiwen Li +5
Deep neural networks for semantic segmentation rely on large-scale annotated datasets, leading to an annotation bottleneck that motivates few shot semantic segmentation (FSS) which…
Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
JiaKui Hu, Zhengjian Yao, Lujia Jin +2
Translation equivariance is a fundamental inductive bias in image restoration, ensuring that translated inputs produce translated outputs. Attention mechanisms in modern restoratio…