6 papers
RADAR: Redundancy-Aware Diffusion for Multi-Agent Communication Structure Generation
Zhen Zhang, Wanjing Zhou, Juncheng Li +3
Compared with individual agents, large language model based multi-agent systems have shown great capabilities consistently across diverse tasks, including code generation, mathemat…
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
Pengxin Xu, Xincheng Lin, Luping Xiao +5
General scene perception has progressed from object recognition toward open-vocabulary grounding, part localization, and affordance prediction. Yet these capabilities are often rea…
Orthogonal Spatial-temporal Distributional Transfer for 4D Generation
Wei Liu, Shengqiong Wu, Bobo Li +4
In the AIGC era, generating high-quality 4D content has garnered increasing research attention. Unfortunately, current 4D synthesis research is severely constrained by the lack of…
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
Yanlin Li, Minghui Guo, Kaiwen Zhang +13
In real-world multimodal applications, systems usually need to comprehend arbitrarily combined and interleaved multimodal inputs from users, while also generating outputs in any in…
TAIL: Text-Audio Incremental Learning
Yingfei Sun, Xu Gu, Wei Ji +3
Many studies combine text and audio to capture multi-modal information but they overlook the model's generalization ability on new datasets. Introducing new datasets may affect the…
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
Bobo Li, Yuheng Wang, Hao Fei +4
Online form filling is a common yet labor-intensive task involving extensive keyboard and mouse interactions. Despite the long-standing vision of automating this process with "one…