1 citations · 1 across the 4 of their papers we have counts for
4 papers
Pretrained Image-Text Models are Secretly Video Captioners
Chunhui Zhang, Yiren Jian, Zhongyu Ouyang +1
Developing video captioning models is computationally expensive. The dynamic nature of video also complicates the design of multimodal models that can effectively caption these seq…
Learning Musical Representations for Music Performance Question Answering
Xingjian Diao, Chunhui Zhang, Tingxuan Wu +4
Music performances are representative scenarios for audio-visual modeling. Unlike common scenarios with sparse audio, music performances continuously involve dense audio signals th…
Is It Navajo? Accurate Language Detection in Endangered Athabaskan Languages
Ivory Yang, Weicheng Ma, Chunhui Zhang +1
Endangered languages, such as Navajo - the most widely spoken Native American language - are significantly underrepresented in contemporary language technologies, exacerbating the…
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding
Xingjian Diao, Chunhui Zhang, Weiyi Wu +5
Multimodal foundation models (MFMs) have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. However, these models fa…