A Distributed Frank-Wolfe Framework for Learning Low-Rank Matrices with the Trace Norm
arXiv:1712.07495 · doi:10.1007/s10994-018-5713-5
Abstract
We consider the problem of learning a high-dimensional but low-rank matrix from a large-scale dataset distributed over several machines, where low-rankness is enforced by a convex trace norm constraint. We propose DFW-Trace, a distributed Frank-Wolfe algorithm which leverages the low-rank structure of its updates to achieve efficiency in time, memory and communication usage. The step at the heart of DFW-Trace is solved approximately using a distributed version of the power method. We provide a theoretical analysis of the convergence of DFW-Trace, showing that we can ensure sublinear convergence in expectation to an optimal solution with few power iterations per epoch. We implement DFW-Trace in the Apache Spark distributed programming framework and validate the usefulness of our approach on synthetic and real data, including the ImageNet dataset with high-dimensional features extracted from a deep neural network.
References in corpus (11)
- Quantum state tomography via compressed sensing
- A Simpler Approach to Matrix Completion
- Recovering low-rank matrices from few coefficients in any basis
- Global Optimality of Local Search for Low Rank Matrix Recovery
- Consistency of trace norm minimization
- On the Global Linear Convergence of Frank-Wolfe Optimization Variants
- Block-Coordinate Frank-Wolfe Optimization for Structural SVMs
- Decentralized Frank-Wolfe Algorithm for Convex and Non-convex Problems
- Faster Rates for the Frank-Wolfe Method over Strongly-Convex Sets
- Variance-Reduced and Projection-Free Stochastic Optimization
- Parallel and Distributed Block-Coordinate Frank-Wolfe Algorithms
Cited by in corpus (9)
- One Sample Stochastic Frank-Wolfe
- Quantized Frank-Wolfe: Faster Optimization, Lower Communication, and Projection Free
- Privacy-preserving Channel Estimation in Cell-free Hybrid Massive MIMO Systems
- Neural Conditional Gradients
- Communication-Efficient Asynchronous Stochastic Frank-Wolfe over Nuclear-norm Balls
- Scalable Projection-Free Optimization
- Distributed Primal-Dual Optimization for Online Multi-Task Learning
- RC Circuits based Distributed Conditional Gradient Method
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum