• Home
  • Search
  • Parameter reduction in convolutional neural networks with kernel transposition
  • https://doi.org/10.1186/s43067-025-00287-wCopy DOI Icon

Parameter reduction in convolutional neural networks with kernel transposition

Show More
  • Abstract
  • PDF
  • Literature Map
  • References
  • Similar Papers
Abstract

Deep convolutional neural networks (CNNs) are a popular choice for many image classification tasks due to their good performance. In general, there is a correlation between the generalization performance of a deep learning model, the number of training samples available, and the number of free parameters in the model. If there are too many free parameters in the model relative to the amount of training data, the model overfits the training data. Reducing the number of trainable parameters in CNNs while maintaining performance levels is an active area of research. This is important, especially in resource-constrained environments where either the size of the training data is insufficient or the memory of the hardware where the model is to be deployed is limited, such as in mobile vision applications. In this paper, kernel transposition is proposed as a method for the reduction of free parameters in the model’s convolutional layers. This allows kernel reuse, in which a model learns a given kernel only once but uses it twice. The learned kernels and their transposes can either be used in sequence or in parallel to one another. These form series and parallel convolutional modules respectively. The modules are used as replacements of the traditional convolutional layers in a CNN model. The use of the modules reduces the number of free parameters and computational costs associated with convolution operations in the model without necessarily compromising its generalization performance. The proposed method is generic and can be adapted to existing state-of-the-art network architectures as demonstrated in the experiments using established architectures. The proposed method was validated on the CIFAR-10 and CIFAR-100 standard datasets using a series of experiments with five model architectures: a basic 5-convolutional layer CNN model, the ResNet-56, small MobileNetV3, EfficientNetB0 and ConvNext model architectures; all trained from scratch. Models based on standard convolutional layers (for basic CNN and ResNet-56 architectures) and depth-wise convolutional layers (for MobileNetV3, EfficientNetB0 and ConvNext architectures) were compared against those based on the proposed convolutional modules. The models based on the standard and depth-wise convolutional layers were used as baselines. The comparisons were in terms of the number of parameters, computational efficiencies (Floating Point Operations / FLOPs) and classification accuracies of the models. For a kernel size k, the experiments showed that, despite improvements on parameter and computational efficiencies, models based on the proposed modules for the basic CNN (k = 3) architecture had lower accuracy levels compared to their baseline model. However, the huge saving on the number of parameters (> 88%) and FLOPs (> 84%) is significant compared to the drop in accuracy (< 4%). For the MobileNetV3 (k = 3 or k = 5) and EfficientNetB0 (k = 3 or k = 5), only the model based on the series module had lower accuracy than the baseline models. The models based on the parallel module were generally at par (within a margin of $$\pm$$ 1.5%) in terms of accuracy with their respective baselines. In relation to the ConvNext (k = 7) and the ResNet-56 (k = 3 or k = 11, or k = 19) architectures, other than the improved parameter and computational efficiencies, the classification accuracies of all the models based on the proposed modules were either at par or better than those of their baseline models. In all the above architectures, the use of the proposed modules reduced the number of convolutional parameters in the networks by at least 78% with fewer FLOPs relative to the baseline models

Loading PDF

Similar Papers
  • Conference Article
  • Citations7

Twitter Sentiment Analysis About Public Opinion on 4G Smartfren Network Services Using Convolutional Neural Network

  • Oct 01, 2019
  • Muhammad Radifan Aldiansyah +1
  • Book Chapter

Efficient Neural Network Architectures

  • Jan 12, 2022
  • Han Cai +1
  • Research Article
  • Citations10

Optimizing convolutional neural networks on multi-core vector accelerator

  • Jun 08, 2022
  • Parallel Computing
  • Zhong Liu +4
  • Research Article

Long-term prediction of heart failure using an ECG-based fine-tuned self-supervised neural network model identifies 23 genetic loci implicating electrophysiological mechanisms

  • Nov 01, 2025
  • European Heart Journal
  • E Caballero +5
  • Research Article

Hybrid Kolmogorov-Arnold and convolutional neural network model for single-lead electrocardiogram classification

  • Oct 01, 2025
  • TELKOMNIKA (Telecommunication Computing Electronics and Control)
  • Marlin Ramadhan Baidillah +7
  • Conference Article
  • Citations4

CNN-BPSO approach to Select Optimal Values of CNN Parameters for Software Requirements Classification

  • Dec 10, 2020
  • Manjubala Bisi +1
  • Conference Article
  • Citations15

Impact of Variation in Number of Channels in CNN Classification model for Cervical Cancer Detection

  • Sep 03, 2021
  • Nitin Kumar Chauhan +1
  • Research Article
  • Citations2

Recognizing Pneumonia Infection in Chest X-Ray Using Deep Learning

  • Oct 10, 2023
  • Matrik Jurnal Manajemen Teknik Informatika dan Rekayasa Komputer
  • Ni Wayan Sumartini Saraswati +5
  • Research Article
  • Citations10

Cyclic Sparsely Connected Architectures for Compact Deep Convolutional Neural Networks

  • Oct 01, 2021
  • IEEE Transactions on Very Large Scale Integration (VLSI) Systems
  • Morteza Hosseini +6
  • Conference Article
  • Citations6

Early Blight and Late Blight Disease Prediction using CNN for Potato Leaves

  • Sep 08, 2022
  • Katiki Sai Krishna +1
  • Research Article
  • Citations1

Performance Comparison of Convolutional Neural Network and Long Short-Term Memory for the Classification of Handwritten Digits

  • Jan 01, 2025
  • International Journal of Research and Innovation in Applied Science
  • Oluwatobi Joel Toyobo +3
  • PDF
  • Research Article

Classification of Transmission Line Ground Short Circuit Fault Based on Convolutional Neural Network

  • Jan 01, 2022
  • Energy Engineering
  • Tao Guo +5
  • Research Article
  • Citations3

Hybrid Solution Through Systematic Electrical Impedance Tomography Data Reduction and CNN Compression for Efficient Hand Gesture Recognition on Resource-Constrained IoT Devices

  • Feb 14, 2025
  • Future Internet
  • Salwa Sahnoun +5
  • Research Article
  • Citations31

Convolutional support vector machines for speech recognition

  • Dec 11, 2018
  • International Journal of Speech Technology
  • Vishal Passricha +1
  • Research Article
  • Citations140

Understanding the learning mechanism of convolutional neural networks in spectral analysis

  • Apr 08, 2020
  • Analytica Chimica Acta
  • Xiaolei Zhang +8
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.