activity
20232026
most citedDual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization

6 citations · 8 across the 8 of their papers we have counts for

collaborators

9 papers

cs.CL2026

Lngram v2: Latent N-Gram Memory with Interpretable Discrete Representations

Yunao Zheng, Bin Wen, Xiaojie Wang +12

Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent condition…

cs.CL2026

Lngram: N-gram Conditional Memory in Latent Space

Yunao Zheng, Guoyang Xia, Xiaojie Wang +1

Sequence modeling requires both compositional reasoning and local static knowledge retrieval, yet standard Transformers handle both through dense computation. Engram partially deco…

cs.CL2026

ROSA-Tuning: Enhancing Long-Context Modeling via Suffix Matching

Yunao Zheng, Xiaojie Wang, Lei Ren +1

Long-context capability and computational efficiency are among the central challenges facing today's large language models. Existing efficient attention methods reduce computationa…

cs.CV2025

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding

Xiaoyi Bao, Chenwei Xie, Hao Tang +4

In recent years, the introduction of Multi-modal Large Language Models (MLLMs) into video understanding tasks has become increasingly prevalent. However, how to effectively integra…

cs.CV2025

Aligned Better, Listen Better for Audio-Visual Large Language Models

Yuxin Guo, Shuailei Ma, Shijie Ma +7

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large la…

cs.CV2024

CoReS: Orchestrating the Dance of Reasoning and Segmentation

Xiaoyi Bao, Siyang Sun, Shuailei Ma +5

The reasoning segmentation task, which demands a nuanced comprehension of intricate queries to accurately pinpoint object regions, is attracting increasing attention. However, Mult…