activity
20242026
most citedAdversarial Rademacher Complexity of Deep Neural Networks

5 citations · 5 across the 1 of their papers we have counts for

collaborators

6 papers

cs.LG20265 cited

Adversarial Rademacher Complexity of Deep Neural Networks

Jiancong Xiao, Yanbo Fan, Ruoyu Sun +1

Deep neural networks (DNNs) are highly vulnerable to adversarial attacks. Ideally, a robust model should perform well on both perturbed training data and unseen perturbed test data…

cs.LG2026

Understanding Adversarial Imitation Learning in Small Sample Regime: A Stage-coupled Analysis

Tian Xu, Ziniu Li, Yang Yu +1

Imitation learning learns a policy from expert trajectories. While the expert data is believed to be crucial for imitation quality, it was found that a kind of imitation learning a…

cs.LG2025

Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation

Ziniu Li, Congliang Chen, Tianyun Yang +5

Large Language Models (LLMs) can self-improve through reinforcement learning, where they generate trajectories to explore and discover better solutions. However, this exploration p…

cs.LG2025

Preserving Diversity in Supervised Fine-Tuning of Large Language Models

Ziniu Li, Congliang Chen, Tian Xu +4

Large Language Models (LLMs) typically rely on Supervised Fine-Tuning (SFT) to specialize in downstream tasks, with the Cross Entropy (CE) loss being the de facto choice. However,…

cs.LG2025

Adam-mini: Use Fewer Learning Rates To Gain More

Yushun Zhang, Congliang Chen, Ziniu Li +6

We propose Adam-mini, an optimizer that achieves on par or better performance than AdamW with 50% less memory footprint. Adam-mini reduces memory by cutting down the learning rate…

cs.LG2024

Exploring the Generalization Capabilities of AID-based Bi-level Optimization

Congliang Chen, Li Shen, Zhiqiang Xu +3

Bi-level optimization has achieved considerable success in contemporary machine learning applications, especially for given proper hyperparameters. However, due to the two-level op…