• Home
  • Search
  • Dual stochastic natural gradient descent
  • https://doi.org/10.1007/s41884-025-00165-4Copy DOI Icon

Dual stochastic natural gradient descent

Show More
  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • References
  • Similar Papers
Abstract

The multinomial logistic regression (MLR) model is widely used in statistics and machine learning. On the one hand, stochastic gradient descent (SGD) is the most common approach for determining the parameters of a such model in big data scenarios, due to its simplicity and low computational complexity property. Furthermore, SGD has proven convergence under reasonable conditions. However, SGD has slow sub-linear rates of convergence and it often reduces convergence speed due to the plateau phenomenon. On the other hand, stochastic natural gradient descent (SNGD), proposed by Amari, is a manifold optimization method shown to be Fisher efficient when it converges, but its convergence properties remain unproven and it is often computationally prohibitive for models with a large number of parameters. Here, we propose dual stochastic natural gradient descent (DSNGD), a stochastic optimization method for MLR based on manifold optimization concepts. In the discrete scenario, DSNGD (i) has linear per-iteration computational complexity in the number of parameters, and (ii) is proven to converge. To achieve (i) we leverage the dual flatness of the family of joint distributions for MLR to simplify computations. To ensure (ii) DSNGD builds on the foundational ideas of convergent stochastic natural gradient descent (CSNGD), a variant of SNGD with guaranteed convergence, using an independent sequence to construct a bounded approximation of the natural gradient. By generalizing a result from Sunehag et al., we prove that DSNGD converges in the discrete case and maintains linear computational complexity per iteration. Beyond its convergence property and linear computational complexity, DSNGD empirically demonstrates fast convergence comparable to SNGD, improves upon SGD performance, and exhibits stability where SNGD does not.

Loading PDF

Similar Papers
  • Video Transcripts

E5 SGD with Low-Dimensional Gradients with Applications to Private and Distributed Learning

  • Jul 17, 2021
  • Underline Science Inc.
  • Shiva Kasiviswanathan
  • Conference Article
  • Citations1

Evaluating the performance of hyperparameters for unbiased and fair machine learning

  • Apr 02, 2024
  • Vy Bui +4
  • Video Transcripts

An Efficient Empirical Solver for Localized Multiple Kernel Learning Via DNNs

  • Dec 29, 2020
  • Underline Science Inc.
  • Ziming Zhang
  • Conference Article
  • Citations2

Low-Cost Lipschitz-Independent Adaptive Importance Sampling of Stochastic Gradients

  • Jan 10, 2021
  • Huikang Liu +3
  • Research Article
  • Citations118

Preconditioned Stochastic Gradient Descent.

  • Mar 09, 2017
  • IEEE Transactions on Neural Networks and Learning Systems
  • Xi-Lin Li
  • Research Article
  • Citations5

Adaptive proximal SGD based on new estimating sequences for sparser ERM

  • Apr 19, 2023
  • Information Sciences
  • Zhuan Zhang +1
  • Book Chapter

Gradient‐Based Optimizers for Statistics and Machine Learning

  • Feb 15, 2022
  • Wiley StatsRef: Statistics Reference Online
  • Cho‐Jui Hsieh
  • Conference Article
  • Citations17

Stochastic Gradient Descent on Modern Hardware: Multi-core CPU or GPU? Synchronous or Asynchronous?

  • May 01, 2019
  • Yujing Ma +2
  • Research Article
  • Citations3

Convergence Analysis of Accelerated Stochastic Gradient Descent Under the Growth Condition

  • Dec 06, 2023
  • Mathematics of Operations Research
  • You-Lin Chen +2
  • Research Article
  • Citations78

Decoupled stochastic parallel gradient descent optimization for adaptive optics: integrated approach for wave-front sensor information fusion.

  • Feb 01, 2002
  • Journal of the Optical Society of America A
  • Mikhail A Vorontsov
  • Research Article
  • Citations10

Risk optimization using the Chernoff bound and stochastic gradient descent

  • Apr 09, 2022
  • Reliability Engineering & System Safety
  • André Gustavo Carlon +4
  • Conference Article
  • Citations1

Transient Failures Detection in Data Transfer for Wireless Sensors based Communication using Stochastic Gradient Descent with Momentum

  • Jan 20, 2022
  • S Jayapratha +2
  • Conference Article
  • Citations2

A Chaos Theory Approach to Understand Neural Network Optimization

  • Nov 01, 2021
  • Michele Sasdelli +3
  • Book Chapter

CHEAPS2AGA: Bounding Space Usage in Variance-Reduced Stochastic Gradient Descent over Streaming Data and Its Asynchronous Parallel Variants

  • Jan 01, 2020
  • Yaqiong Peng +4
  • Video Transcripts

Sample-based Approximation of Nash in Large Many-Player Games via Gradient Descent

  • Apr 20, 2022
  • Underline Science Inc.
  • Janos Kramar +8
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.