most citedTake the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization

3 citations · 3 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG2024

Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis

Hongru Yang, Bhavya Kailkhura, Zhangyang Wang +1

Understanding the training dynamics of transformers is important to explain the impressive capabilities behind large language models. In this work, we study the dynamics of trainin…

cs.LG20243 cited

Take the Bull by the Horns: Hard Sample-Reweighted Continual Training Improves LLM Generalization

Xuxi Chen, Zhendong Wang, Daouda Sow +5

In the rapidly advancing arena of large language models (LLMs), a key challenge is to enhance their capabilities amid a looming shortage of high-quality training data. Our study st…

cs.LG2023

Rethinking PGD Attack: Is Sign Function Necessary?

Junjie Yang, Tianlong Chen, Xuxi Chen +2

Neural networks have demonstrated success in various domains, yet their performance can be significantly degraded by even a small input perturbation. Consequently, the construction…

cs.CV2023

Meta ControlNet: Enhancing Task Adaptation via Meta Learning

Junjie Yang, Jinze Zhao, Peihao Wang +2

Diffusion-based image synthesis has attracted extensive attention recently. In particular, ControlNet that uses image-based prompts exhibits powerful capability in image tasks such…

cs.CV2023

Data Distillation Can Be Like Vodka: Distilling More Times For Better Quality

Xuxi Chen, Yu Yang, Zhangyang Wang +1

Dataset distillation aims to minimize the time and memory needed for training deep networks on large datasets, by creating a small set of synthetic images that has a similar genera…