Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
FlexControl: Computation-Aware ControlNet with Differentiable Router for Text-to-Image Generation
Zheng Fang, Lichuan Xiang, Xu Cai +2
ControlNet offers a powerful way to guide diffusion-based generative models, yet most implementations rely on ad-hoc heuristics to choose which network blocks to control-an approac…
cs.LG2024
No More Adam: Learning Rate Scaling at Initialization is All You Need
Minghao Xu, Lichuan Xiang, Xu Cai +1
In this work, we question the necessity of adaptive gradient methods for training deep neural networks. SGD-SaI is a simple yet effective enhancement to stochastic gradient descent…
cs.LG2023
How Much Is Hidden in the NAS Benchmarks? Few-Shot Adaptation of a NAS Predictor
Hrushikesh Loya, Łukasz Dudziak, Abhinav Mehrotra +4
Neural architecture search has proven to be a powerful approach to designing and refining neural networks, often boosting their performance and efficiency over manually-designed va…