activity
20242026
collaborators

25 papers

eess.AS2026

BAMU: Bitstream-Aware Marginal-Utility Allocation for Frozen Pretrained Neural Speech Codecs

Mingyu Zhao, Zijian Lin, Yutang Feng +6

Pretrained neural speech codecs typically use a fixed residual vector quantization (RVQ) depth for all frames, ignoring temporal variation in quantization difficulty. We propose BA…

cs.CV2026

RAVA: Retrieval-Augmented Viewpoint Alignment for Subject-Driven Image Generation

Qiwei Yan, Zhiqiang Yuan, Chongyang Li +4

Reference-driven image generation has made rapid progress on identity preservation, but reliable viewpoint control across different subjects remains poorly understood. The difficul…

cs.AI2026

PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums

Qiwei Yan, Zhiqiang Yuan, Zexi Jia +4

Longitudinal personal albums are weak-schema multimodal databases: noisy perceptual records whose key facts require joins across faces, text, timestamps, locations, and repeated ev…

cs.CL2026

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues

Hanbo Bi, Zhiqiang Yuan, Chongyang Li +7

With the widespread adoption of multi-modal communication platforms, long-form dialogues interleaving text and images have become increasingly common. Users often need to retrieve…

cs.SD2026

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling

Xiaoyue Duan, Nanxing Hu, Yutang Feng +4

Recent song generation systems can synthesize realistic audio, yet generating complete songs remains challenging for two reasons. First, explicit song-level arrangement planning re…

cs.CL2026

PhotoCraft: Agentic Reasoning with Hierarchical Self-Evolving Memory for Deep Image Search

Kailin Lyu, Zhiqiang Yuan, Jianwei He +9

Deep Image Search requires multi-step reasoning over rich contextual cues, such as time, location, and event relations. However, most existing LLM-based agents are stateless and re…