• Home
  • Search
  • Convex-Concave Programming: An Effective Alternative for Optimizing Shallow Neural Networks
  • https://doi.org/10.1109/tetci.2024.3502463Copy DOI Icon

Convex-Concave Programming: An Effective Alternative for Optimizing Shallow Neural Networks

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

In this study, we address the challenges of non-convex optimization in neural networks (NNs) by formulating the training of multilayer perceptron (MLP) NNs as a difference of convex functions (DC) problem. Utilizing the basic convex–concave algorithm to solve our DC problems, we introduce two alternative optimization techniques, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DC-GD</i> and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DC-OPT</i>, for determining MLP parameters. By leveraging the non-uniqueness property of the convex components in DC functions, we generate strongly convex components for the DC NN cost function. This strong convexity enables our proposed algorithms, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DC-GD</i> and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DC-OPT</i>, to achieve an <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">iteration complexity</i> of <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$O\left(\log \left(\frac{1}{\varepsilon }\right)\right)$</tex-math></inline-formula>, surpassing that of other solvers, such as stochastic gradient descent (<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">SGD</i>), which has an <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">iteration complexity</i> of <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$O\left(\frac{1}{\varepsilon }\right)$</tex-math></inline-formula>. This improvement raises the convergence rate from sublinear (<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">SGD</i>) to linear (ours) while maintaining comparable <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">total computational costs</i>. Furthermore, conventional NN optimizers like <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">SGD</i>, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">RMSprop</i>, and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Adam</i> are highly sensitive to the learning rate, adding computational overhead for practitioners in selecting an appropriate learning rate. In contrast, our <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DC-OPT</i> algorithm is hyperparameter-free (i.e., it requires no learning rate), and our <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">DC-GD</i> algorithm is less sensitive to the learning rate, offering comparable accuracy to other solvers. Additionally, we extend our approach to a convolutional NN architecture, demonstrating its applicability to modern NNs. We evaluate the performance of our proposed algorithms by comparing them to conventional optimizers such as <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">SGD</i>, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">RMSprop</i>, and <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Adam</i> across various test cases. The results suggest that our approach is a viable alternative for optimizing shallow MLP NNs.

Similar Papers
  • Supplementary Content
  • Citations15

Text Complexity Analysis of Chinese and foreign academic English writing via mobile devices based on neural network and deep learning

  • May 17, 2022
  • Library Hi Tech
  • Qiucheng Liu
  • Research Article
  • Citations16

How Deep Learning Networks could be Designed to Locate Mineral Deposits

  • Apr 01, 2021
  • Journal of Earth Science
  • Donald A Singer
  • PDF
  • Research Article
  • Citations27

Peptide-Major Histocompatibility Complex Class I Binding Prediction Based on Deep Learning With Novel Feature.

  • Nov 28, 2019
  • Frontiers in genetics
  • Tianyi Zhao +3
  • Research Article

Wasserstein Distributionally Robust Shallow Convex Neural Networks

  • Aug 26, 2025
  • INFORMS Journal on Optimization
  • Julien Pallage +1
  • PDF
  • Research Article
  • Citations7

RazorNet: Adversarial Training and Noise Training on a Deep Neural Network Fooled by a Shallow Neural Network

  • Jul 23, 2019
  • Big Data and Cognitive Computing
  • Shayan Taheri +2
  • Research Article

Enhanced flow rate prediction of disturbed pipe flow using a shallow neural network

  • Jan 01, 2025
  • Flow
  • Christoph Wilms +4
  • Research Article
  • Citations15

A Robust DNS Flood Attack Detection with a Hybrid Deeper Learning Model

  • Mar 24, 2022
  • Computers and Electrical Engineering
  • Ömer Kasim
  • Conference Article
  • Citations2

Classification of Urine Odour Using Machine Learning Methods

  • May 29, 2022
  • Yuxin Xing +1
  • Research Article
  • Citations1

Geospatial Feature-Based Path Loss Prediction at 1800 MHz in Covenant University Campus with Tree Ensembles, Kernel-Based Methods, and a Shallow Neural Network

  • Oct 20, 2025
  • Electronics
  • Marta Moreno-Cuevas +4
  • Research Article
  • Citations121

Forecasting of solar and wind power using LSTM RNN for load frequency control in isolated microgrid

  • Jun 30, 2020
  • International Journal of Modelling and Simulation
  • Dhananjay Kumar +3
  • Conference Article
  • Citations35

SAR ATR based on dividing CNN into CAE and SNN

  • Sep 01, 2015
  • Xuan Li +4
  • PDF
  • Research Article
  • Citations5

A Low-Resolution Used Electronic Parts Image Dataset for Sorting Application

  • Jan 14, 2023
  • Data
  • Praneel Chand
  • Conference Article
  • Citations14

Enabling Incremental Knowledge Transfer for Object Detection at the Edge

  • Jun 01, 2020
  • Mohammad Farhadi +3
  • Research Article

Shallow Neural Networks and Laplace's Equation on the Half-Space with Dirichlet Boundary Data

  • Dec 06, 2024
  • Pittsburgh Interdisciplinary Mathematics Review
  • Malhar Vaishampayan
  • PDF
  • Research Article
  • Citations1

RETRACTED ARTICLE: Data simulation of optimal model for numerical solution of differential equations based on deep learning and genetic algorithm

  • May 06, 2023
  • Soft Computing
  • Li Jing
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.