activity
20232026
most citedMega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

16 citations · 27 across the 31 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal

Hanting Wang, Shengpeng Ji, Shulei Wang +4

Image restoration under adverse weather conditions has been extensively explored, leading to numerous high-performance methods. In particular, recent advances in All-in-One approac…

cs.CV2025

Open-set Cross Modal Generalization via Multimodal Unified Representation

Hai Huang, Yan Xia, Shulei Wang +6

This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates…

cs.CV2025

IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

Hanting Wang, Tao Jin, Wang Lin +4

Bridge models in image restoration construct a diffusion process from degraded to clear images. However, existing methods typically require training a bridge model from scratch for…

cs.CV2025

Astrea: A MOE-based Visual Understanding Model with Progressive Alignment

Xiaoda Yang, JunYu Lu, Hongshun Qiu +12

Vision-Language Models (VLMs) based on Mixture-of-Experts (MoE) architectures have emerged as a pivotal paradigm in multimodal understanding, offering a powerful framework for inte…

cs.CV2024

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Weiheng Lu, Jian Li, An Yu +3

Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain cha…