47 citations · 54 across the 5 of their papers we have counts for
5 papers
DiffAgent: Fast and Accurate Text-to-Image API Selection with Large Language Model
Lirui Zhao, Yue Yang, Kaipeng Zhang +5
Text-to-image (T2I) generative models have attracted significant attention and found extensive applications within and beyond academic research. For example, the Civitai community,…
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
Junting Chen, Yao Mu, Qiaojun Yu +11
Rapid progress in high-level task planning and code generation for open-world robot manipulation has been witnessed in Embodied AI. However, previous studies put much effort into g…
Towards Unified and Effective Domain Generalization
Yiyuan Zhang, Kaixiong Gong, Xiaohan Ding +4
We propose , a novel and fied framework for omain eneralization that is capable of significantly enhancing the out-of-distribu…
Meta-Transformer: A Unified Framework for Multimodal Learning
Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang +4
Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging…
DiffRate : Differentiable Compression Rate for Efficient Vision Transformers
Mengzhao Chen, Wenqi Shao, Peng Xu +6
Token compression aims to speed up large-scale vision transformers (e.g. ViTs) by pruning (dropping) or merging tokens. It is an important but challenging task. Although recent adv…