- Research Article
14
- 10.1016/j.neunet.2023.12.027
Neural networks with ReLU powers need less depth
- Dec 19, 2023
- Neural Networks
- Kurt Izak M Cabanilla + 2 more +2
Neural networks with ReLU powers need less depth
In deep learning, the choice of activation function plays a vital role in enhancing model performance. We propose AHerfReLU, a novel activation function that combines the rectified linear unit (ReLU) function with the error function (erf), complemented by a regularization term 1/(1 + x 2 ), ensuring smooth gradients even for negative inputs. The function is zero centered, bounded below, and nonmonotonic, offering significant advantages over traditional activation functions like ReLU. We compare AHerfReLU with 10 adaptive activation functions and state‐of‐the‐art activation functions, including ReLU, Swish, and Mish. Experimental results show that replacing ReLU with AHerfReLU leads to 3.18% improvement in Top‐1 accuracy on the LeNet network for the CIFAR100 dataset, 0.63% improvement on CIFAR10%, and 1.3% improvement in mean average precision (mAP) on the SSD300 model in the Pascal VOC dataset. Our results demonstrate that AHerfReLU enhances model performance, offering improved accuracy, loss reduction, and convergence stability. The function outperforms existing activation functions, providing a promising alternative for deep learning tasks.
Neural networks with ReLU powers need less depth
Neural networks with ReLU powers need less depth
RMAF: Relu-Memristor-Like Activation Function for Deep Learning
Activation functions facilitate deep neural networks by introducing non-linearity to the learning process. The non-linearity feature gives the neural network the ability to learn complex patterns. Recently, the most widely used activation function is the Rectified Linear Unit (ReLU). Though, other various existing activation including hand-designed alternatives to ReLU have been proposed. However, none has succeeded in replacing ReLU due to their existing inconsistencies. In this work, activation function called ReLUMemristor-like Activation Function (RMAF) is proposed to leverage benefits of negative values in neural networks. RMAF introduces a constant parameter (α) and a threshold parameter (p) making the function smooth, non-monotonous, and introduces non-linearity in the network. Our experiments show that, the RMAF works better than ReLU and other activation functions on deeper models and across number of challenging datasets. Firstly, experiments are performed by training and classifying on multi-layer perceptron (MLP) over benchmark data such as the Wisconsin breast cancer, MNIST, Iris and Car evaluation. RMAF achieves high performance of 98.74%, 99.67%, 98.81% and 99.42% respectively, compared to Sigmoid, Tanh and ReLU. Secondly, experiments were performed on convolution neural network (ResNet) over MNIST, CIFAR-10 and CIFAR-100 data and observed the proposed activation function achieves higher performance accuracy of 99.73%, 98.77% and 79.82% respectively than Tanh, ReLU and Swish. Additionally, we experimented our work on deep networks i.e. squeeze network (SqueezeNet), Dense connected neural network (DenseNet121) and ImageNet dataset, which RMAF produced the best performance. We note that, the RMAF converges faster than the other functions and can replace ReLU in any neural network due to the efficiency, scalability and its similarity to both ReLU and Swish.
Read moreEvolutionary optimization of deep learning activation functions
The choice of activation function can have a large effect on the performance of a neural network. While there have been some attempts to hand-engineer novel activation functions, the Rectified Linear Unit (ReLU) remains the most commonly-used in practice. This paper shows that evolutionary algorithms can discover novel activation functions that outperform ReLU. A tree-based search space of candidate activation functions is defined and explored with mutation, crossover, and exhaustive search. Experiments on training wide residual networks on the CIFAR-10 and CIFAR-100 image datasets show that this approach is effective. Replacing ReLU with evolved activation functions results in statistically significant increases in network accuracy. Optimal performance is achieved when evolution is allowed to customize activation functions to a particular task; however, these novel activation functions are shown to generalize, achieving high performance across tasks. Evolutionary optimization of activation functions is therefore a promising new dimension of metalearning in neural networks.
Read moreDeep Leaky Single-peaked Triangle Neural Networks
Recently, Deep learning has made a great deal of success in processing images, audios, and natural languages and so on. The activation function is one of the key factors in Deep learning. In this paper, according to characteristics of biological neurons, an improved Leaky Single-Peaked Triangle Linear Unit (LSPTLU) activation function is presented for the right-hand response unbounded of Rectified Linear Unit (ReLU) and Leaky ReLU (LReLU). LSPTLU is more in line with the biological neuron essence and achieves the excellent performance of equivalent or beyond ReLU and LReLU on different datsets, e.g., MNIST, Fashion-MNIST, SVHN, IMAGENET, CALTECH101 and CIFAR10 datasets.
Read moreOptimizing Deep Learning Models with Custom ReLU for Breast Cancer Histopathology Image Classification
Purpose: The prompt identification of breast cancer is crucial in preventing the considerable damage inflicted by this dangerous form of cancer, which is widely distributed across the globe. This study seeks to refine the efficacy of a deep learning-driven approach for the precise diagnosis of breast cancer by employing diverse bespoke Rectified Linear Units (ReLU) to improve the model's performance and reduce inaccuracies within the system. Method: This study focuses on analyzing a deep learning approach utilizing the BreakHis dataset with 7,909 images, incorporating changes to the ReLU activation function across different pre-trained CNN models. It then evaluates performance through measures such as accuracy, precision, recall, and F1-Score. Result: Based on our experiment results, show that the DenseNet201 models with a custom LeakyReLU excel beyond the typical ReLU, achieving the highest accuracy, recall, and F1-Score at 99.21%, 99.21%, and 99.11%, respectively. Simultaneously, ResNet152, utilizing LessNegativeReLU ( ), achieved the highest precision at 99.11%. The VGG11 model exhibited the most notable performance enhancement, with improvements ranging from 1.39% to 1.59%. Novelty: The research is original in optimizing a model for accurate breast cancer diagnosis. The proposed model is superior to the model utilizing the default activation function. This finding indicates that the study significantly enhances performance while effectively minimizing errors, thereby necessitating further exploration into the effectiveness of the customized activation function when applied to other medical imaging modalities.
Read moreAverage biased ReLU based CNN descriptor for improved face retrieval
The convolutional neural networks (CNN), including AlexNet, GoogleNet, VGGNet, etc. extract features for many computer vision problems which are very discriminative. The trained CNN model over one dataset performs reasonably well whereas on another dataset of similar type the hand-designed feature descriptor outperforms the same trained CNN model. The Rectified Linear Unit (ReLU) layer discards some values in order to introduce the non-linearity. In this paper, it is proposed that the discriminative ability of deep image representation using trained model can be improved by Average Biased ReLU (AB-ReLU) at the last few layers. Basically, AB-ReLU improves the discriminative ability in two ways: 1) it exploits some of the discriminative and discarded negative information of ReLU and 2) it also neglects the irrelevant and positive information used in ReLU. The VGGFace model trained in MatConvNet over the VGG-Face dataset is used as the feature descriptor for face retrieval over other face datasets. The proposed approach is tested over six challenging, unconstrained and robust face datasets (PubFig, LFW, PaSC, AR, FERET and ExtYale) and also on a large scale face dataset (PolyUNIR) in retrieval framework. It is observed that the AB-ReLU outperforms the ReLU when used with a pre-trained VGGFace model over the face datasets. The validation error by training the network after replacing all ReLUs with AB-ReLUs is also observed to be favorable over each dataset. The AB-ReLU even outperforms the state-of-the-art activation functions, such as Sigmoid, ReLU, Leaky ReLU and Flexible ReLU over all seven face datasets.
Read moreSoft-Clipping Swish: A Novel Activation Function for Deep Learning
This study aims to contribute to the improvement of the network's performance through developing a novel activation function. Over time, many activation functions have been proposed in order to solve the issues of the previous functions. We note here more than 50 activation functions that have been proposed, some of them being very popular such as sigmoid, Rectified Linear Unit (ReLU), Swish, Mish but not only. The main idea of this study that stays behind our proposal is a simple one, based on a very popular function called Swish, which is a composition function, having in its componence sigmoid function and ReLU function. Starting from this activation function we decided to ignore the negative region in the way the Rectified Linear Unit does but being different than that one mentioned through a nonlinear curve assured by the Swish positive region. The idea has been come up from a current function called Soft Clipping. We tested this proposal on more datasets in Computer Vision on classification tasks showing its high potential, here we mention MNIST, Fashion-MNIST, CIFAR-10, CIFAR-100 using two popular architectures: LeNet-5 and ResNet20 version 1.
Read moreTuning a Fully Convolutional Network for Velocity Model Estimation
Different parameters of a fully convolutional network (FCN) are experimented to evaluate which combination predicts sound velocity models from a single configuration of seismic modeling. The evaluation is made considering some fixed parameters of the deep learning model, such as number of epochs, batch size and loss function, but with variations of the optimizer and activation function. The considered optimizers were RMSprop, Adam and Adamax, whilst the activations functions were the Rectified Linear Unit (ReLU), Leaky ReLU, Exponential Linear Unit (ELU) and Parametric ReLU (PReLU). Five metrics were used to evaluate the model during the testing stage: R2, Pearson's r, factor of two, mean absolute error and mean squared error. To the extent of these experiments, it was found that the optimizers have much more influence than the activation functions when determining the resolution of the output model. The best combination was the one using the PReLU activation function with the Adamax optimizer.
Read moreActivation Function Optimizations for Capsule Networks
Classical Convolutional Neural Networks, or ConvNets, have been the benchmarks for most object classification and face recognition tasks despite suffering from major limitations such as the inability to capture spatial co-locality between data points and favoring invariance over equivariance. Hence, a hierarchical routing layered architecture called Capsule Networks was proposed to overcome shortcomings of ConvNets. Capsules replace average or max pooling techniques of ConvNets with dynamic routing abilities between lower level and higher level neural units which better capture hierarchical relationships within the data and introduced reconstruction regularization mechanisms which deals with equivariance properties. By overcoming existing limitations, Capsules have proven themselves to be potential benchmarks in object segmentation, detection and reconstruction. Capsules have achieved state of the art results on the fundamental MNIST (Modified National Institute of Standards and Technology) handwritten digit dataset by reducing the ConvNets test error benchmark of 0.39% to 0.25%. In order to further augment this distinction, we experimented with five activation units such as sigmoid, e-Swish, Swish, variants of Rectified Linear Units (ReLU) like Parametric ReLU (PReLU), leaky ReLU (lReLU) and Scaled Exponential Linear Units (SELU), on two fundamental datasets - MNIST and cifar10. Based on these experimental results, we establish that e-Swish, and ReLU variants better optimize the Capsule architecture as compared to the currently used ReLU activation function.
Read moreAdaptive Rectified Linear Unit (Arelu) for Classification Problems to Solve Dying Problem in Deep Learning
A convolutional neural network (CNN) is a subset of machine learning as well as one of the different types of artificial neural networks that are used for different applications and data types. Activation functions (AFs) are used in this type of network to determine whether or not its neurons are activated. One non-linear AF named as Rectified Linear Units (ReLU) which involves a simple mathematical operations and it gives better performance. It avoids rectifying vanishing gradient problem that inherents older AFs like tanh and sigmoid. Additionally, it has less computational cost. Despite these advantages, it suffers from a problem called Dying problem. Several modifications have been appeared to address this problem, for example; Leaky ReLU (LReLU). The main concept of our algorithm is to improve the current LReLU activation functions in mitigating the dying problem on deep learning by using the readjustment of values (changing and decreasing value) of the loss function or cost function while number of epochs are increased. The model was trained on the MNIST dataset with 20 epochs and achieved lowest misclassification rate by 1.2%. While optimizing our proposed methods, we received comparatively better results in terms of simplicity, low computational cost, and with no hyperparameters.
Read moreBeyond ReLU: Unlocking Superior Plant Disease Recognition with Swish
Plant disease identification plays a crucial role in agricultural management, but the traditional methods are time-consuming and imprecise. This work contributes to the ongoing efforts to develop reliable and efficient solutions for automated plant disease diagnosis, ultimately aiding in the timely management and mitigation of agricultural challenges. This study investigated the effectiveness of various deep convolutional neural networks (CNNs) for automated plant disease identification from leaf images. By exploring deep and transfer learning techniques such as CNNs, InceptionNet, DenseNet 121, and ResNet-50, a Deep Convolutional Neural Network (DCNN) is proposed to categorize the leaf disease. Different activation functions, specifically Rectified Linear Unit (ReLU) and Swish, are employed to investigate their impact on model performance. The experiments revealed that the DCNN architecture, when paired with the Swish activation function, demonstrated 96% accuracy in plant disease identification.
Read moreEvolution of activation functions for deep learning-based image classification
Activation functions (AFs) play a pivotal role in the performance of neural networks. The Rectified Linear Unit (ReLU) is currently the most commonly used AF. Several replacements to ReLU have been suggested but improvements have proven inconsistent. Some AFs exhibit better performance for specific tasks, but it is hard to know a priori how to select the appropriate one(s). Studying both standard fully connected neural networks (FCNs) and convolutional neural networks (CNNs), we propose a novel, three-population, coevolutionary algorithm to evolve AFs, and compare it to four other methods, both evolutionary and non-evolutionary. Tested on four datasets -- MNIST, FashionMNIST, KMNIST, and USPS -- coevolution proves to be a performant algorithm for finding good AFs and AF architectures.
Read moreYolk color measurement using image processing and deep learning
A high yellow yolk color of laying hens is required by customer. As yolk color measurement is determined by visual perception, color score may be expressed differently. The objective of this study was to develop the recognition of yolk color using red green blue (RGB) image and deep learning. The three hundred and fifty-three RGB images were obtained. The rectified linear unit (ReLU) and softmax were used as the activation function. An optimizer was configured with Adam, and categorical crossentropy was used as a loss function. The results showed that the loss had decreased to 0.45 and 0.63, whereas the accuracy had increased and reached 0.80 and 0.76 for training dataset and testing dataset, respectively. For evaluation, the loss value was 0.27 and 0.63, whereas the accuracy value was 0.90 and 0.76 for training dataset and testing dataset, respectively. The average f1-score was 0.76, whereas the highest precision (1.00) was observed in color score 5, 6 and 8. In conclusion, RGB image can be used as an alternative method to classify yolk color score with lower cost of analysis for egg producers in the near future.
Read moreAn improved fault diagnosis method based on deep wavelet neural network
Deep learning has been successfully applied to the field of fault diagnosis in recent years. Due to the advantages of deep belief network (DBN) in fitting nonlinear complex systems and the ability of wavelet analysis in time-frequency analysis, in this paper, an improved fault diagnosis method based on a deep wavelet neural network (DWNN), which combines the DBN with morlet activation functions, is proposed for fault diagnosis of reciprocating compressor. A five-layer DBN using sigmoid, tanh, rectified linear unit (ReLU) and morlet wavelet functions as the activation functions of hidden layers separately is proposed for fault diagnosis of reciprocating compressor. As the contrast, a three-layer back propagation neural network (BPNN) using the same four activation functions separately is proposed for fault diagnosis of reciprocating compressor. The experimental results show that, the fault diagnosis rate of five-layer DBN is higher than the three-layer BPNN. The method based on DWNN can make the fault diagnosis rate reach 100% within short time. Compared with using other activation functions, the DWNN architecture requires less epochs to train the model.
Read moreClassification of Alzheimer's Disease Based on Eight-Layer Convolutional Neural Network with Leaky Rectified Linear Unit and Max Pooling.
Alzheimer's disease (AD) is a progressive brain disease. The goal of this study is to provide a new computer-vision based technique to detect it in an efficient way. The brain-imaging data of 98AD patients and 98 healthy controls was collected using data augmentation method. Then, convolutional neural network (CNN) was used, CNN is the most successful tool in deep learning. An 8-layer CNN was created with optimal structure obtained by experiences. Three activation functions (AFs): sigmoid, rectified linear unit (ReLU), and leaky ReLU. The three pooling-functions were also tested: average pooling, max pooling, and stochastic pooling. The numerical experiments demonstrated that leaky ReLU and max pooling gave the greatest result in terms of performance. It achieved a sensitivity of 97.96%, a specificity of 97.35%, and an accuracy of 97.65%, respectively. In addition, the proposed approach was compared with eight state-of-the-art approaches. The method increased the classification accuracy by approximately 5% compared to state-of-the-art methods.
Read more