papers

Publications (24)

cs.LG2018

Block-Normalized Gradient Method: An Empirical Study for Training Deep Neural Network

Adams Wei Yu, Lei Huang, Qihang Lin +2

In this paper, we propose a generic and simple strategy for utilizing stochastic gradient information in optimization. The technique essentially contains two consecutive steps in e…

cs.LG2017

Orthogonal Weight Normalization: Solution to Optimization over Multiple Dependent Stiefel Manifolds in Deep Neural Networks

Lei Huang, Xianglong Liu, Bo Lang +3

Orthogonal matrix has shown advantages in training Recurrent Neural Networks (RNNs), but such matrix is limited to be square for the hidden-to-hidden transformation in RNNs. In thi…

stat.ML2016

An Improved Gap-Dependency Analysis of the Noisy Power Method

Maria Florina Balcan, Simon S. Du, Yining Wang +1

We consider the noisy power method algorithm, which has wide applications in machine learning and statistics, especially those related to principal component analysis (PCA) under r…

eess.SY2015

Efficient Structured Matrix Rank Minimization

Adams Wei Yu, Wanli Ma, Yaoliang Yu +2

We study the problem of finding structured low-rank matrices using nuclear norm regularization where the structure is encoded by a linear map. In contrast to most known approaches…

cs.CL2024

Large Language Models Cannot Self-Correct Reasoning Yet

Jie Huang, Xinyun Chen, Swaroop Mishra +4

Large Language Models (LLMs) have emerged as a groundbreaking technology with their unparalleled text generation capabilities across various applications. Nevertheless, concerns pe…

cs.CV2022

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Zirui Wang, Jiahui Yu, Adams Wei Yu +3

With recent progress in joint modeling of visual and textual representations, Vision-Language Pretraining (VLP) has achieved impressive performance on many multimodal downstream ta…