11 citations · 11 across the 8 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Monkey King Bang: A Unified Scientific Multimodal Foundation Model
Hesen Chen, Xinyu Su, Xiaomeng Yang +11
Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition. Existing systems are either specia…
cs.LG2023
Data-Juicer: A One-Stop Data Processing System for Large Language Models
Daoyuan Chen, Yilun Huang, Zhijian Ma +10
The immense evolution in Large Language Models (LLMs) has underscored the importance of massive, heterogeneous, and high-quality data. A data recipe is a mixture of data from diffe…
cs.LG2021
Fine-Grained AutoAugmentation for Multi-Label Classification
Ya Wang, Hesen Chen, Fangyi Zhang +4
Data augmentation is a commonly used approach to improving the generalization of deep learning models. Recent works show that learned data augmentation policies can achieve better…