• Home
  • Search
  • Nonparametric Perspective of Deep Learning
  • https://doi.org/10.25394/pgs.13360247.v1Copy DOI Icon

Nonparametric Perspective of Deep Learning

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Models built with deep neural network (DNN) can handle complicated real-world data extremely well, seemingly without suffering from the curse of dimensionality or the non-convex optimization. To contribute to the theoretical understanding of deep learning, this work studies the nonparametric perspective of DNNs by considering the following questions: (1) What is the underlying estimation problem and what are the most appropriate data assumptions? (2) What is the corresponding optimal convergence rate and does the curse of dimensionality occur? (3) Is the optimal rate achievable for DNN estimators and is there any optimization guarantee? These questions are investigated on two of the most fundamental problems --- regression and classification. Specifically, statistical optimality of DNN estimators is established under various settings with special focuses on the curse of dimensionality and optimization guarantee.In the classic binary classification problem, statistical optimal convergence rates that suffer less from the curse of dimensionality are established under two settings:(1) Under the smooth boundary assumption, I show that DNN classifiers with proper architectures can benefit from the compositional smoothness structure underlying the high dimensional data in the sense that the optimal convergence rates only depend on some effective dimension d*, potentially much smaller than the data dimension d. (2) Under a novel teacher-student framework that assumes the Bayes classifier to be expressed as ReLU neural networks, I obtain a dimension-free rate of convergence O(n^{-2/3}) for DNN classifiers, which is also proven optimal. The optimization of DNN is highly complicated with generally no algorithmic guarantee on finding the global minimizer. To this end, I turn to the recently proposed neural tangent kernel (NTK) literature where the similarity between overparametrized DNN trained by gradient descent (GD) and kernel methods are established. Specifically, through a comprehensive analysis of L2-regularized GD trajectories, I prove that for overparametrized one-hidden-layer ReLU neural networks with L2 regularization, the output from GD is close to that from the kernel ridge regression with the corresponding NTK and optimal rate of L2 estimation error can be achieved.

Similar Papers
  • PDF
  • Research Article

Optimal deep neural network architecture design with improved generalization for data-driven cooling load estimation problem

  • May 02, 2025
  • Neural Computing and Applications
  • Baris Baykant Alagoz +4
  • Book Chapter

Unboundedness of Linear Regions of Deep ReLU Neural Networks

  • Jan 01, 2022
  • Anton Ponomarchuk +2
  • Research Article
  • Citations41

Kernel ridge vs. principal component regression: Minimax bounds and the qualification of regularization operators

  • Jan 01, 2017
  • Electronic Journal of Statistics
  • Lee H Dicker +2
  • Conference Article
  • Citations14

Joint Optimization of Quantization and Structured Sparsity for Compressed Deep Neural Networks

  • May 01, 2019
  • Gaurav Srivastava +5
  • Conference Article
  • Citations36

ReForm

  • Jun 02, 2019
  • Zirui Xu +3
  • Research Article
  • Citations2

E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFS

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Ziyang Zhang +4
  • Research Article
  • Citations12

Optimizing DNN training with pipeline model parallelism for enhanced performance in embedded systems

  • Apr 06, 2024
  • Journal of Parallel and Distributed Computing
  • Md Al Maruf +3
  • Research Article
  • Citations1

Kernel ridge regression improving based on golden eagle optimization algorithm for multi-class classification

  • Jul 25, 2025
  • Statistics, Optimization & Information Computing
  • Shaimaa Mahmood +1
  • Research Article
  • Citations1

Deep ReLU neural networks overcome the curse of dimensionality when approximating semilinear partial integro-differential equations

  • Mar 26, 2025
  • Analysis and Applications
  • Ariel Neufeld +2
  • Research Article
  • Citations46

Generalization error of random feature and kernel methods: Hypercontractivity and kernel matrix concentration

  • Jul 01, 2022
  • Applied and Computational Harmonic Analysis
  • Song Mei +2
  • Conference Article
  • Citations13

Survey on Comparative Study of Pruning Mechanism on MobileNetV3 Model

  • Jun 25, 2021
  • Shiva V Naik +7
  • Research Article
  • Citations15

Sketch Kernel Ridge Regression Using Circulant Matrix: Algorithm and Theory.

  • Oct 29, 2019
  • IEEE transactions on neural networks and learning systems
  • Rong Yin +3
  • Research Article
  • Citations39

Fine-Grained Detection of Driver Distraction Based on Neural Architecture Search

  • Sep 01, 2021
  • IEEE Transactions on Intelligent Transportation Systems
  • Jie Chen +6
  • Research Article
  • Citations2

On rate-optimal nonparametric wavelet regression with long memory moving average errors

  • Jun 06, 2013
  • Statistical Inference for Stochastic Processes
  • Linyuan Li +1
  • Supplementary Content

In Theory and Practice - On the Rate of Convergence of Implementable Neural Network Regression Estimates

  • Jan 01, 2021
  • TUbilio (Technical University of Darmstadt)
  • Alina Braun
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.