most citedMeta-Transformer: A Unified Framework for Multimodal Learning

47 citations · 54 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2024

DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model

Lirui Zhao, Yue Yang, Kaipeng Zhang +5

Text-to-image (T2I) generative models have attracted significant attention and found extensive applications within and beyond academic research. For example, the Civitai community,…

cs.RO20243 cited

RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation

Junting Chen, Yao Mu, Qiaojun Yu +11

Rapid progress in high-level task planning and code generation for open-world robot manipulation has been witnessed in Embodied AI. However, previous studies put much effort into g…

cs.CV20231 cited

Towards Unified and Effective Domain Generalization

Yiyuan Zhang, Kaixiong Gong, Xiaohan Ding +4

We propose , a novel and fied framework for omain eneralization that is capable of significantly enhancing the out-of-distribu…

cs.CV202347 cited

Meta-Transformer: A Unified Framework for Multimodal Learning

Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang +4

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging…

cs.CV20233 cited

DiffRate : Differentiable Compression Rate for Efficient Vision Transformers

Mengzhao Chen, Wenqi Shao, Peng Xu +6

Token compression aims to speed up large-scale vision transformers (e.g. ViTs) by pruning (dropping) or merging tokens. It is an important but challenging task. Although recent adv…