Geometry, Computation, and Optimality in Stochastic Optimization

September 23, 2019 · Declared Dead · 🏛 NeurIPS 2019

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Chen Cheng, Daniel Levy, John C. Duchi arXiv ID 1909.10455 Category math.OC: Optimization & Control Cross-listed cs.IT, cs.LG, stat.ML Citations 11 Venue NeurIPS 2019 Last Checked 4 months ago

Abstract

We study computational and statistical consequences of problem geometry in stochastic and online optimization. By focusing on constraint set and gradient geometry, we characterize the problem families for which stochastic- and adaptive-gradient methods are (minimax) optimal and, conversely, when nonlinear updates -- such as those mirror descent employs -- are necessary for optimal convergence. When the constraint set is quadratically convex, diagonally pre-conditioned stochastic gradient methods are minimax optimal. We provide quantitative converses showing that the ``distance'' of the underlying constraints from quadratic convexity determines the sub-optimality of subgradient methods. These results apply, for example, to any $\ell_p$-ball for $p < 2$, and the computation/accuracy tradeoffs they demonstrate exhibit a striking analogy to those in Gaussian sequence models.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Optimization & Control

R.I.P. 👻 Ghosted

Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent

Xiangru Lian, Ce Zhang, ... (+4 more)

math.OC 🏛 NeurIPS 📚 1.4K cites 9 years ago

R.I.P. 👻 Ghosted

Local SGD Converges Fast and Communicates Little

Sebastian U. Stich

math.OC 🏛 ICLR 📚 1.2K cites 8 years ago

R.I.P. 👻 Ghosted

On Lazy Training in Differentiable Programming

Lenaic Chizat, Edouard Oyallon, Francis Bach

math.OC 🏛 NeurIPS 📚 930 cites 7 years ago

📚 📚 The Cartographer

A Review on Bilevel Optimization: From Classical to Evolutionary Approaches and Applications

Ankur Sinha, Pekka Malo, Kalyanmoy Deb

math.OC 🏛 IEEE TEC 📚 840 cites 9 years ago

R.I.P. 👻 Ghosted

Learned Primal-dual Reconstruction

Jonas Adler, Ozan Öktem

math.OC 🏛 IEEE TMI 📚 834 cites 8 years ago

R.I.P. 👻 Ghosted

On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport

Lenaic Chizat, Francis Bach

math.OC 🏛 NeurIPS 📚 805 cites 8 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Federated Learning: Strategies for Improving Communication Efficiency

Jakub Konečný, H. Brendan McMahan, ... (+4 more)

cs.LG 🏛 arXiv 📚 5.2K cites 9 years ago

R.I.P. 👻 Ghosted

In-Datacenter Performance Analysis of a Tensor Processing Unit

Norman P. Jouppi, Cliff Young, ... (+73 more)

cs.AR 🏛 ISCA 📚 5.1K cites 9 years ago

R.I.P. 👻 Ghosted

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Hoo-Chang Shin, Holger R. Roth, ... (+7 more)

cs.CV 🏛 IEEE TMI 📚 4.9K cites 10 years ago

R.I.P. 👻 Ghosted

Explanation in Artificial Intelligence: Insights from the Social Sciences

Tim Miller

cs.AI 🏛 AI 📚 4.9K cites 9 years ago