collaborators

9 papers

cs.CV2025

Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis

Hao Tang, Ling Shao, Zhenyu Zhang +2

We propose a novel spatial-temporal graph Mamba (STG-Mamba) for the music-guided dance video synthesis task, i.e., to translate the input music to a dance video. STG-Mamba consists…

cs.CV2025

Replace in Translation: Boost Concept Alignment in Counterfactual Text-to-Image

Sifan Li, Ming Tao, Hao Zhao +2

Text-to-Image (T2I) has been prevalent in recent years, with most common condition tasks having been optimized nicely. Besides, counterfactual Text-to-Image is obstructing us from…

cs.CV2025

SAMba-UNet: SAM2-Mamba UNet for Cardiac MRI in Medical Robotic Perception

Guohao Huo, Ruiting Dai, Ling Shao +1

To address complex pathological feature extraction in automated cardiac MRI segmentation, we propose SAMba-UNet, a novel dual-encoder architecture that synergistically combines the…

cs.CV2025

MambaIC: State Space Models for High-Performance Learned Image Compression

Fanhu Zeng, Hao Tang, Yihua Shao +3

A high-performance image compression algorithm is crucial for real-time information transmission across numerous fields. Despite rapid progress in image compression, computational…

cs.CV2025

Frequency Domain Enhanced U-Net for Low-Frequency Information-Rich Image Segmentation in Surgical and Deep-Sea Exploration Robots

Guohao Huo, Ruiting Dai, Jinliang Liu +2

In deep-sea exploration and surgical robotics scenarios, environmental lighting and device resolution limitations often cause high-frequency feature attenuation. Addressing the dif…

cs.CV2025

Enhanced Multi-Scale Cross-Attention for Person Image Generation

Hao Tang, Ling Shao, Nicu Sebe +1

In this paper, we propose a novel cross-attention-based generative adversarial network (GAN) for the challenging person image generation task. Cross-attention is a novel and intuit…