4 papers · 1 filter
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
Yujin Tang, Chenming Shang, Ruize Xu +1
Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to rem…
HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
Ling Yang, Xinchen Zhang, Ye Tian +4
The remarkable success of the autoregressive paradigm has made significant advancement in Multimodal Large Language Models (MLLMs), with powerful models like Show-o, Transfusion an…
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance
Zhao Wang, Hao Wen, Lingting Zhu +3
Character video generation is a significant real-world application focused on producing high-quality videos featuring specific characters. Recent advancements have introduced vario…
Understanding Multimodal Deep Neural Networks: A Concept Selection View
Chenming Shang, Hengyuan Zhang, Hao Wen +1
The multimodal deep neural networks, represented by CLIP, have generated rich downstream applications owing to their excellent performance, thus making understanding the decision-m…