activity
20242026
collaborators

20 papers

cs.CV2026

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

Dongxu Ge, Shansong Liu, Cheng Gong +3

As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted incre…

cs.AI2026

HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics

Shufan Jiang, Sizhou Chen, Chios Chen +3

Creating an immersive and interactive theatrical experience is a long-term goal in the field of interactive narrative. The emergence of large language models (LLMs) provides a new…

cs.RO2026

VTouch++: A Multimodal Dataset with Vision-Based Tactile Enhancement for Bimanual Manipulation

Qianxi Hua, Xinyue Li, Zheng Yan +4

Embodied intelligence has advanced rapidly in recent years; however, bimanual manipulation-especially in contact-rich tasks remains challenging. This is largely due to the lack of…

eess.AS2026

High-Fidelity Generative Audio Compression at 0.275kbps

Hao Ma, Ruihao Jing, Shansong Liu +4

High-fidelity general audio compression at ultra-low bitrates is crucial for applications ranging from low-bandwidth communication to generative audio-language modeling. Traditiona…

cs.CV2026

Loupe: A Generalizable and Adaptive Framework for Image Forgery Detection

Yuchu Jiang, Jiaming Chu, Jian Zhao +5

The proliferation of generative models has raised serious concerns about visual content forgery. Existing deepfake detection methods primarily target either image-level classificat…

eess.AS2025

Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models

Ruihao Jing, Cheng Gong, Yu Jiang +5

Rare words remain a critical bottleneck for speech-to-text systems. While direct fine-tuning improves recognition of target words, it often incurs high cost, catastrophic forgettin…