- Research Article
193
- 10.1016/s0968-090x(03)00020-2
Incident detection using support vector machines
- Jun 01, 2003
- Transportation Research Part C: Emerging Technologies
- Fang Yuan + 1 more +1
Incident detection using support vector machines
Support vector machine (SVM) is a new sort of machine learning method based on Structure Risk Minimization (SRM) principle, which has high generalization capability. Many problems with small samples, nonlinearity or high dimension in pattern recognition could be solved by the method. In this paper, the traffic data on freeway were taken as research objects and an information fusion algorithm based on SVM about freeway incident detection was proposed. A SVM was trained and tested using the data obtained from the simulation under the condition of incident and non-incident. Compared with the multi-layer feed forward neural network (MLF) algorithm trained with the same data, the simulation results showed that the SVM offers a lower misclassification rate, higher correct detection rate and lower false alarm, and it can improve the detection performance.
Incident detection using support vector machines
Incident detection using support vector machines
Comparing the performance of support vector machines to regression with structural risk minimisation
The structural risk minimisation (SRM) principle based on the statistical learning theory of Vapnik aims to prevent the phenomenon of overfitting by balancing the complexity of models with their fit to the data. This principle has been embodied in support vector machines, a widely acclaimed generic approach to machine learning. This paper investigates the performance of the SRM principle in its application to standard least-squares regression and compares it with its integration with support vector machines.
Read moreApplication of Support Vector Machine to Predict 5-year Survival Status of Patients with Nasopharyngeal Carcinoma after Treatment
Objective: Support Vector Machine (SVM) is a machine-learning method, based on the principle of structural risk minimization, which performs well when applied to data outside the training set. In this paper, SVM was applied to predict 5-year survival status of patients with nasopharyngeal carcinoma (NPC) after treatment, we expect to find a new way for prognosis studies in cancer so as to assist right clinical decision for individual patient. Methods: Two modelling methods were used in the study; SVM network and a standard parametric logistic regression were used to model 5-year survival status. And the two methods were compared on a prospective set of patients not used in model construction via receiver operating characteristic (ROC) curve analysis. Results: The SVM1, trained with the 25 original input variables without screening, yielded a ROC area of 0.868, at sensitivity to mortality of 79.2% and the specificity of 94.5%. Similarly, the SVM2, trained with 9 input variables which were obtained by optimal input variable selection from the 25 original variables by logistic regression screening, yielded a ROC area of 0.874, at a sensitivity to mortality of 79.2% and the specificity of 95.6%, while the logistic regression yielded a ROC area of 0.751 at a sensitivity to mortality of 66.7% and gave a specificity of 83.5%. Conclusion: SVM found a strong pattern in the database predictive of 5-year survival status. The logistic regression produces somewhat similar, but better, results. These results show that the SVM models have the potential to predict individual patient's 5-year survival status after treatment, and to assist the clinicians for making a good clinical decision.
Read moreComparison Support Vector Machine and Fuzzy Possibilistic C-Means based on the kernel for Knee Osteoarthritis data Classification
Osteoarthritis is a chronic joint disease that occurs when the protective cartilage that cushions the ends of bones wears down over time and fails to be repaired. The common form of the disease is knee osteoarthritis while it can affect all body parts with joints, such as hands, ankles, hips, and spine. The major cause of knee osteoarthritis is the continuous depletion of its cartilage. During the diagnosis, machine learning is used because early prevention is necessary for proper treatment. This study, therefore, considers classification methods of Support Vector Machine (SVM) and clustering methods using fuzzy clusterings such as Fuzzy C-Means (FCM), Fuzzy Possibilistic C-Means (FPCM), and Fuzzy Possibilistic C-Means based on kernel (FPCMK) to analyze of knee osteoarthritis. SVM is a machine learning technique that works based on the principle of structural risk minimization (SRM) to obtain the best hyperplane to separate two or more classes in input space. Otherwise, the fuzzy clustering is to determine the value of a distance and to know and measure the similarity of each object to be observed. FPCMK uses the kernel Radial Base Function (RBF) in the fuzzy clustering method. The kernel function is applicable for handling non-separable data problems. This method will be compared to the level of the measured parameter; their accuracy, recall, precision, and f1 score. The greatest level of accuracy is generated from SVM with an accuracy value of 86.7%, then followed by FPCMK with an accuracy value of 85.5%.
Read moreEnsemble Learning with Support Vector Machines for Bond Rating
Bond rating is regarded as an important event for measuring financial risk of companies and for determining the investment returns of investors. As a result, it has been a popular research topic for researchers to predict companies' credit ratings by applying statistical and machine learning techniques. The statistical techniques, including multiple regression, multiple discriminant analysis (MDA), logistic models (LOGIT), and probit analysis, have been traditionally used in bond rating. However, one major drawback is that it should be based on strict assumptions. Such strict assumptions include linearity, normality, independence among predictor variables and pre-existing functional forms relating the criterion variablesand the predictor variables. Those strict assumptions of traditional statistics have limited their application to the real world. Machine learning techniques also used in bond rating prediction models include decision trees (DT), neural networks (NN), and Support Vector Machine (SVM). Especially, SVM is recognized as a new and promising classification and regression analysis method. SVM learns a separating hyperplane that can maximize the margin between two categories. SVM is simple enough to be analyzed mathematical, and leads to high performance in practical applications. SVM implements the structuralrisk minimization principle and searches to minimize an upper bound of the generalization error. In addition, the solution of SVM may be a global optimum and thus, overfitting is unlikely to occur with SVM. In addition, SVM does not require too many data sample for training since it builds prediction models by only using some representative sample near the boundaries called support vectors. A number of experimental researches have indicated that SVM has been successfully applied in a variety of pattern recognition fields. However, there are three major drawbacks that can be potential causes for degrading SVM's performance. First, SVM is originally proposed for solving binary-class classification problems. Methods for combining SVMs for multi-class classification such as One-Against-One, One-Against-All have been proposed, but they do not improve the performance in multi-class classification problem as much as SVM for binary-class classification. Second, approximation algorithms (e.g. decomposition methods, sequential minimal optimization algorithm) could be used for effective multi-class computation to reduce computation time, but it could deteriorate classification performance. Third, the difficulty in multi-class prediction problems is in data imbalance problem that can occur when the number of instances in one class greatly outnumbers the number of instances in the other class. Such data sets often cause a default classifier to be built due to skewed boundary and thus the reduction in the classification accuracy of such a classifier. SVM ensemble learning is one of machine learning methods to cope with the above drawbacks. Ensemble learning is a method for improving the performance of classification and prediction algorithms. AdaBoost is one of the widely used ensemble learning techniques. It constructs a composite classifier by sequentially training classifiers while increasing weight on the misclassified observations through iterations. The observations that are incorrectly predicted by previous classifiers are chosen more often than examples that are correctly predicted. Thus Boosting attempts to produce new classifiers that are better able to predict examples for which the current ensemble's performance is poor. In this way, it can reinforce the training of the misclassified observations of the minority class. This paper proposes a multiclass Geometric Mean-based Boosting (MGM-Boost) to resolve multiclass prediction problem. Since MGM-Boost introduces the notion of geometric mean into AdaBoost, it can perform learning process considering the geometric mean-based accuracy and errors of multiclass. This study applies MGM-Boost to the real-world bond rating case for Korean companies to examine the feasibility of MGM-Boost. 10-fold cross validations for threetimes with different random seeds are performed in order to ensure that the comparison among three different classifiers does not happen by chance. For each of 10-fold cross validation, the entire data set is first partitioned into tenequal-sized sets, and then each set is in turn used as the test set while the classifier trains on the other nine sets. That is, cross-validated folds have been tested independently of each algorithm. Through these steps, we have obtained the results for classifiers on each of the 30 experiments. In the comparison of arithmetic mean-based prediction accuracy between individual classifiers, MGM-Boost (52.95%) shows higher prediction accuracy than both AdaBoost (51.69%) and SVM (49.47%). MGM-Boost (28.12%) also shows the higher prediction accuracy than AdaBoost (24.65%) and SVM (15.42%)in terms of geometric mean-based prediction accuracy. T-test is used to examine whether the performance of each classifiers for 30 folds is significantly different. The results indicate that performance of MGM-Boost is significantly different from AdaBoost and SVM classifiers at 1% level. These results mean that MGM-Boost can provide robust and stable solutions to multi-classproblems such as bond rating.
Read moreSLIT: Designing Complexity Penalty for Classification and Regression Trees Using the SRM Principle
The statistical learning theory has formulated the Structural Risk Minimization (SRM) principle, based upon the functional form of risk bound on the generalization performance of a learning machine. This paper addresses the application of this formula, which is equivalent to a complexity penalty, to model selection tasks for decision trees, whereas the quantization of the machine capacity for decision trees is estimated using an empirical approach. Experimental results show that, for either classification or regression problems, this novel strategy of decision tree pruning performs better than alternative methods. We name classification and regression trees pruned by virtue of this methodology as Statistical Learning Intelligent Trees (SLIT).
Read morePREDICTION OF PASSENGER FLOW ON THE HIGHWAY BASED ON THE LEAST SQUARE SUPPOERT VECTOR MACHINE / MAŽIAUSIŲ KVADRATŲ ATRAMINIŲ VEKTORIŲ METODO TAIKYMAS KELEIVIŲ SRAUTUI GREITKELYJE PROGNOZUOTI / ПРИМЕНЕНИЕ МЕТОДА ОПОРНЫХ ВЕКТОРОВ С КВАДРАТИЧНОЙ ФУНКЦИЕЙ ПОТЕРЬ ДЛЯ ПРОГНОЗИРОВАНИЯ ПАССАЖИРСКИХ ПОТОКОВ НА АВТОМАГИСТРАЛЯХ
A support vector machine is a machine learning method based on the statistical learning theory and structural risk minimization. The support vector machine is a much better method than ever, because it may solve some actual problems in small samples, high dimension, nonlinear and local minima etc. The article utilizes the theory and method of support vector machine (SVM) regression and establishes the regressive model based on the least square support vector machine (LS-SVM). Through predicting passenger flow on Hangzhou highway in 2000–2008, the paper shows that the regressive model of LS-SVM has much higher accuracy and reliability of prediction, and therefore may effectively predict passenger flow on the highway. Santrauka Atraminių vektorių metodas (Support Vector Machine – SVM) yra skaičiuojamasis metodas, paremtas statistikos teorija, struktūriniu požiūriu mažinant riziką. SVM metodas, palyginti su kitais metodais, yra patikimesnis metodas, nes juo remiantis galima išspręsti realias problemas, esant įvairioms sąlygoms. Tyrimams naudojama SVM metodo regresijos teorija ir sukuriamas regresinis modelis, kuris grindžiamas mažiausių kvadratų atraminių vektorių metodu (Least Squares Support Vector Machine – LS-SVM). Straipsnio autoriai prognozuoja keleivių srautą Hangdžou (Kinija) greitkelyje 2000–2008 m. Gauti rezultatai rodo, kad regresinis LS-SVM modelis yra labai tikslus ir patikimas, todėl gali būti efektyviai taikomas keleivių srautams prognozuoti greitkeliuose. Резюме Метод опорных векторов (Support Vector Machine – SVM) – это набор аналогичных алгоритмов вида «обучение с учителем», использующихся для задач классификации и регрессионного анализа. Метод SVM принадлежит к семейству линейных классификаторов. Основная идея метода SVM заключается в переводе исходных векторов в пространство более высокой размерности и поиске разделяющей гиперплоскости с максимальным зазором в этом пространстве. Алгоритм работает в предположении, что чем больше разница или расстояние между параллельными гиперплоскостями, тем меньше будет средняя ошибка классификатора. В сравнении с другими методами метод SVM более надежен и позволяет решать проблемы с различными условиями. Для исследования был использован метод SVM и регрессионный анализ, затем создана регрессионная модель, основанная на методе опорных векторов с квадратичной функцией потерь (Least Squares Support Vector Machine – LS-SVM). Авторы прогнозировали пассажирский поток на автомагистрали Ханчжоу (Китай) в 2000–2008 гг. Полученные результаты показывают, что регрессионная модель LS-SVM является надежной и может быть применена для прогнозирования пассажирских потоков на других магистралях.
Read moreOn Improvement on Generalization Performance of Classifier by Using Empirical Risk
A combination classification algorithm, ER-SVM, is proposed to improve the generalization performance of support vector machine (SVM) by directly making full use of the empirical risk (ER) information of SVM in the paper. SVM classification is the implementation of structure risk minimization (SRM) principle. SVM may achieve SRM from the minimal summation of ER and VC confidence according to the theory of VC dimension. However, the ER is seldom zero for a trained SVM in practice. That is, though the minimal summation of ER and VC confidence can be achieved in theory, it is very time-consuming in parameters selection for a given task to make ER zero. In order to overcome such difficulty, a combination classification algorithm is proposed to improve the performance by utilizing ER information. The SR arising from the existing ER is reduced by using aided nearest neighbor method. In addition, the proposed algorithm is independent of training parameters in SVM. The experimental results verify the effectiveness of the proposed algorithm
Read moreSupervised Learning for Visual Pattern Classification
This chapter presents an overview of the topics and major ideas of supervised learning for visual pattern classification. Two prevalent algorithms, i.e., the support vector machine (SVM) and the boosting algorithm, are briefly introduced. SVMs and boosting algorithms are two hot topics of recent research in supervised learning. SVMs improve the generalization of the learning machine by implementing the rule of structural risk minimization (SRM). It exhibits good generalization even when little training data are available for machine training. The boosting algorithm can boost a weak classifier to a strong classifier by means of the so-called classifier combination. This algorithm provides a general way for producing a classifier with high generalization capability from a great number of weak classifiers.
Read moreEfficient optimization of a Ka-Band MMIC sub-harmonically pumped image rejection diode mixer
An efficient optimization technique, support vector regression (SVR) approach, is proposed for designing of Ka-band MMIC sub-harmonically pumped image rejection diode mixer. This SVR approach comes from the support vector machine (SVM) learning theory, which is based on the structural risk minimization (SRM) principle and leads good generalization ability. With this method, a Ka-band MMIC 4th harmonic image rejection diode mixer is designed using comercial United Monolithic Semiconductors process. Details of design approach and outcome of performance simulations is presented.
Read moreA Comparative Study on Machine Learning Techniques for Prediction of Success of Dental Implants
The market demand for dental implants is growing at a significant pace. In practice, some dental implants do not succeed. Important questions in this regard concern whether machine learning techniques could be used to predict whether an implant will be successful and which are the best techniques for this problem. This paper presents a comparative study on machine learning techniques for prediction of success of dental implants. The techniques compared here are: (a) constructive RBF neural networks (RBF-DDA), (b) support vector machines (SVM), (c) k nearest neighbors (kNN), and (d) a recently proposed technique, called NNSRM, which is based on kNN and the principle of structural risk minimization. We present a number of simulations using real-world data. The simulations were carried out using 10-fold cross-validation and the results show that the methods achieve comparable performance, yet NNSRM and RBF-DDA produced smaller classifiers.
Read moreA Learning Machine Approach for Predicting Thermal Comfort Indices
Human thermal comfort is influenced by psychological as well as physiological factors. Several comfort indices, such as PMV, PPD, TSENS, ET*, DISC, and SET* (see nomenclature) have been developed. These indices attempt to correlate human thermal comfort with environmental conditions. This paper describes the use of a learning algorithm “support vector machine (SVM) learning” for prediction of the thermal comfort indices. The SVM is an artificial intelligent approach that can capture the input/output mapping from the given data. Support vector machines were developed based on the Structural Risk Minimization principle. Different sets of representative experimental environmental factors that affect a homogenous person’s thermal balance were used for training the SVM algorithm. The results demonstrate good correlation between SVM predicted values and those obtained from conventional thermal comfort, such as Fanger Model and “2-Node” model. The “trained SVM” with representative data could be easily and more effectively used to predict the indices compared to other conventional estimation methods.
Read moreGlobal Optimization of Support Vector Machines Using Genetic Algorithms for Bankruptcy Prediction
One of the most important research issues in finance is building accurate corporate bankruptcy prediction models since they are essential for the risk management of financial institutions. Thus, researchers have applied various data-driven approaches to enhance prediction performance including statistical and artificial intelligence techniques. Recently, support vector machines (SVMs) are becoming popular because they use a risk function consisting of the empirical error and a regularized term which is derived from the structural risk minimization principle. In addition, they don't require huge training samples and have little possibility of overfitting. However, in order to use SVM, a user should determine several factors such as the parameters of a kernel function, appropriate feature subset, and proper instance subset by heuristics, which hinders accurate prediction results when using SVM. In this study, we propose a novel approach to enhance the prediction performance of SVM for the prediction of financial distress. Our suggestion is the simultaneous optimization of the feature selection and the instance selection as well as the parameters of a kernel function for SVM by using genetic algorithms (GAs). We apply our model to a real-world case. Experimental results show that the prediction accuracy of conventional SVM may be improved significantly by using our model.
Read moreECG Signal Classification Method Based on Structural Risk Minimization
Arrhythmia stands as a primary contributor to cardiovascular disease-associated mortality. Therefore, the classification and monitoring of abnormal electrocardiogram (ECG) signals are of paramount importance for preventive purposes. Although deep - learning - based ECG classification methods have yielded promising outcomes, they frequently encounter challenges in optimizing performance across diverse patient datasets. To overcome these limitations, this research endeavors to enhance the generalization ability of deep - learning models for ECG signal classification. It achieves this by integrating structural risk minimization principles and incorporating RR interval information into the classification process. A convolutional neural network (CNN) founded on structural risk minimization is proposed. Instead of employing the traditional cross-entropy loss, this study adopts a loss function inspired by support vector machine (SVM) classifiers to optimize the CNN. Moreover, the RR interval information, which is often lost during beat segmentation, is manually extracted and integrated into the CNN network to improve classification accuracy. The proposed method attains an accuracy, specificity, and sensitivity of 88.2% respectively, demonstrating superior performance when compared to traditional and existing methods. This improvement underscores the efficacy of the structural risk minimization approach and the integration of RR interval information in enhancing the model's generalization across patient datasets. The method's convenience and effectiveness render it particularly well-suited for real-time application in wearable devices, facilitating the early detection of abnormal ECG patterns and potentially preventing cardiovascular disease-related fatalities.
Read moreMachinery condition prediction based on wavelet and support vector machine
This paper studies the use of wavelet and support vector machine (SVM) in machinery condition prediction. SVM is based on the VC dimension theory of statistical learning and the principle of structural risk minimization, and has shown advantages in solving the problem with limited sample, nonlinear and high dimensional pattern recognition. The soft failure of mechanical equipment makes its performance drop gradually, which occupies a large proportion and has certain regularity. The performance can be evaluated and predicted through early state monitoring and data analysis. The paper models the vibration signal from the rear pad of a gas blower and analyzes the 1-step and multi-step forecasting of wavelet transformation and SVM (WT-SVM model) and SVM model.
Read more