9 citations · 10 across the 7 of their papers we have counts for
4 papers · 1 filter
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
Yujin Tang, Chenming Shang, Ruize Xu +1
Agent benchmarks for measuring memory largely study textual cases, in which information is deliberately extracted from the environment, written down, and then later retrieved. In o…
HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
Ling Yang, Xinchen Zhang, Ye Tian +4
The remarkable success of the autoregressive paradigm has made significant advancement in Multimodal Large Language Models (MLLMs), with powerful models like Show-o, Transfusion an…
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance
Zhao Wang, Hao Wen, Lingting Zhu +3
Character video generation is a significant real-world application focused on producing high-quality videos featuring specific characters. Recent advancements have introduced vario…
Understanding Multimodal Deep Neural Networks: A Concept Selection View
Chenming Shang, Hengyuan Zhang, Hao Wen +1
The multimodal deep neural networks, represented by CLIP, have generated rich downstream applications owing to their excellent performance, thus making understanding the decision-m…