12 citations · 14 across the 3 of their papers we have counts for
3 papers
cs.MM2024★ 12 cited
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
Zeyu Jin, Jia Jia, Qixin Wang +5
Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elab…
cs.CV2023
Multi-granularity Interaction Simulation for Unsupervised Interactive Segmentation
Kehan Li, Yian Zhao, Zhennan Wang +6
Interactive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and med…
cs.CV2023★ 2 cited
Parallel Vertex Diffusion for Unified Visual Grounding
Zesen Cheng, Kehan Li, Peng Jin +4
Unified visual grounding pursues a simple and generic technical route to leverage multi-task data with less task-specific design. The most advanced methods typically present boxes…