activity
20242026
collaborators

8 papers

cs.CV2026

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

Bingxuan Li, Yiming Cui, Yicheng He +4

Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite recent advancement in video-text-…

cs.HC2026

CoLyricist: Enhancing Lyric Writing with AI through Workflow-Aligned Support

Masahiro Yoshida, Bingxuan Li, Songyan Zhao +4

We propose CoLyricist, an AI-assisted lyric writing tool designed to support the typical workflows of experienced lyricists and enhance their creative efficiency. While lyricists h…

cs.CV2025

METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling

Bingxuan Li, Yiwei Wang, Jiuxiang Gu +2

Chart generation aims to generate code to produce charts satisfying the desired visual properties, e.g., texts, layout, color, and type. It has great potential to empower the autom…

cs.AI2025

Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence

Yining Hong, Rui Sun, Bingxuan Li +7

AI agents today are mostly siloed - they either retrieve and reason over vast amount of digital information and knowledge obtained online; or interact with the physical world throu…

cs.CV2025

Contrastive Visual Data Augmentation

Yu Zhou, Bingxuan Li, Mohan Tang +6

Large multimodal models (LMMs) often struggle to recognize novel concepts, as they rely on pre-trained knowledge and have limited ability to capture subtle visual details. Domain-s…

cs.CL2025

REFFLY: Melody-Constrained Lyrics Editing Model

Songyan Zhao, Bingxuan Li, Yufei Tian +1

Automatic melody-to-lyric (M2L) generation aims to create lyrics that align with a given melody. While most previous approaches generate lyrics from scratch, revision, editing plai…