papers
Publications (3)
cs.CV2026
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
Zhuang Yu, Lei Shen, Jing Zhao +1
Recent multimodal large language models (MLLMs) have shown remarkable progress across vision, audio, and language tasks, yet their performance on long-form, knowledge-intensive, an…
cs.CV2025
Combining Self-attention and Dilation Convolutional for Semantic Segmentation of Coal Maceral Groups
Zhenghao Xi, Zhengnan Lv, Yang Zheng +5
The segmentation of coal maceral groups can be described as a semantic segmentation process of coal maceral group images, which is of great significance for studying the chemical p…
cs.CL2025
Memory Reviving, Continuing Learning and Beyond: Evaluation of Pre-trained Encoders and Decoders for Multimodal Machine Translation
Zhuang Yu, Shiliang Sun, Jing Zhao +2
Multimodal Machine Translation (MMT) aims to improve translation quality by leveraging auxiliary modalities such as images alongside textual input. While recent advances in large-s…