- Research Article
8
- 10.1016/s1874-1029(11)60208-5
Elastic Multiple Kernel Learning
- Jun 01, 2011
- Acta Automatica Sinica
- Zheng-Peng Wu + 1 more +1
Elastic Multiple Kernel Learning
We study support vector machines (SVM) for which the kernel matrix is not specified exactly and it is only known to belong to a given uncertainty set. We consider uncertainties that arise from two sources: (i) data measurement uncertainty, which stems from the statistical errors of input samples; (ii) kernel combination uncertainty, which stems from the weight of individual kernel that needs to be optimized in multiple kernel learning (MKL) problem. Much work has been studied, such as uncertainty sets that allow the corresponding SVMs to be reformulated as semi-definite programs (SDPs), which is very computationally expensive however. Our focus in this paper is to identify uncertainty sets that allow the corresponding SVMs to be reformulated as second-order cone programs (SOCPs), since both the worst case complexity and practical computational effort required to solve SOCPs is at least an order of magnitude less than that needed to solve SDPs of comparable size. In the main part of the paper we propose four uncertainty sets that meet this criterion. Experimental results are presented to confirm the validity of these SOCP reformulations.
Elastic Multiple Kernel Learning
Elastic Multiple Kernel Learning
An efficient multiple-kernel learning for pattern classification
An efficient multiple-kernel learning for pattern classification
Multiple Kernel Learning-Based Uncertainty Set Construction for Robust optimization
In robust optimization (RO), a focal point is the design of uncertainty set that delineates possible realizations of uncertainty since it heavily impacts the robustness of solutions. We propose in this paper a multiple kernel learning (MKL) based support vector clustering (SVC) method for polytopic uncertainty set construction in data-driven RO. By assuming a set of candidate piecewise linear kernel functions, the MKL framework not only derives an enclosing sphere in the input space, but also automatically derives optimal coefficients of kernel functions by only solving a quadratically constrained quadratic program. The learnt sphere turns out to be a compact polyhedral uncertainty set to be used in RO, which helps reducing the conservatism of robust solutions. Meanwhile, although massive data samples and kernel functions are used in MKL, the induced polytopic uncertainty set tends to have a succinct expression, thereby well preserving the tractability of the induced optimization problem. It also allows a decisionmaker to conveniently adjust the conservatism of the data-driven uncertainty set by manipulating only one parameter, which is user-friendly in practice. Numerical case studies are carried out to demonstrate the potential advantages of the proposed method in promoting the practicability of RO techniques.
Read moreSecond-Order Cone Relaxations for Binary Quadratic Polynomial Programs
Several types of relaxations for binary quadratic polynomial programs can be obtained using linear, second-order cone, or semidefinite techniques. In this paper, we propose a general framework to construct conic relaxations for binary quadratic polynomial programs based on polynomial programming. Using our framework, we re-derive previous relaxation schemes and provide new ones. In particular, we present three relaxations for binary quadratic polynomial programs. The first two relaxations, based on second-order cone and semidefinite programming, represent a significant improvement over previous practical relaxations for several classes of nonconvex binary quadratic polynomial problems. From a practical point of view, due to the computational cost, semidefinite-based relaxations for binary quadratic polynomial problems can be used only to solve small to midsize instances. To improve the computational efficiency for solving such problems, we propose a third relaxation based purely on second-order cone programming. Computational tests on different classes of nonconvex binary quadratic polynomial problems, including quadratic knapsack problems, show that the second-order-cone-based relaxation outperforms the semidefinite-based relaxations that are proposed in the literature in terms of computational efficiency, and it is comparable in terms of bounds.
Read moreNew Advances in Convex Optimization and Control Applications
New Advances in Convex Optimization and Control Applications
Kernelized learning in deep scattering convolution networks
This paper addresses the problem of automatic scattering feature selection for signal classification. While features derived from group invariant scattering networks are quite effective for signal classification. We argue that scattering networks are not always the appropriate choice as they are not learned for the objective at hand. In this paper, we explore jointly learning a deep scattering convolution network with a support vector machine by casting the problem as a multiple kernel learning problem. The convolution paths of the network are kernelized respectively to be selected in a large-margin context. We deduce scattering paths from the corresponding kernels after solving the kernel learning problem. Experiments on several datasets demonstrate the effectiveness of the proposed method over state-of-the-art techniques.
Read moreSemidefinite relaxation method for unified near-Field and far-Field localization by AOA
Semidefinite relaxation method for unified near-Field and far-Field localization by AOA
Algorithms for sparse and low-rank optimization: convergence, complexity and applications
Solving optimization problems with sparse or low-rank optimal solutions has been an important topic since the recent emergence of compressed sensing and its matrix extensions such as the matrix rank minimization and robust principal component analysis problems. Compressed sensing enables one to recover a signal or image with fewer observations than the “length” of the signal or image, and thus provides potential breakthroughs in applications where data acquisition is costly. However, the potential impact of compressed sensing cannot be realized without efficient optimization algorithms that can handle extremely large-scale and dense data from real applications. Although the convex relaxations of these problems can be reformulated as either linear programming, second-order cone programming or semidefinite programming problems, the standard methods for solving these relaxations are not applicable because the problems are usually of huge size and contain dense data. In this dissertation, we give efficient algorithms for solving these “sparse” optimization problems and analyze the convergence and iteration complexity properties of these algorithms. Chapter 2 presents algorithms for solving the linearly constrained matrix rank minimization problem. The tightest convex relaxation of this problem is the linearly constrained nuclear norm minimization. Although the latter can be cast and solved as a semidefinite programming problem, such an approach is computationally expensive when the matrices are large. In Chapter 2, we propose fixed-point and Bregman iterative algorithms for solving the nuclear norm minimization problem and prove convergence of the first of these algorithms. By using a homotopy approach together with an approximate singular value decomposition procedure, we get a very fast, robust and powerful algorithm, which we call FPCA (Fixed Point Continuation with Approximate SVD), that can solve very large matrix rank minimization problems. Our numerical results on randomly generated and real matrix completion problems demonstrate that this algorithm is much faster and provides much better recoverability than semidefinite programming solvers such as SDPT3. For example, our algorithm can recover 1000 × 1000 matrices of rank 50 with a relative error of 10−5 in about 3 minutes by sampling only 20 percent of the elements. We know of no other method that achieves as good recoverability. Numerical experiments on online recommendation, DNA microarray data set and image inpainting problems demonstrate the effectiveness of our algorithms. In Chapter 3, we study the convergence/recoverability properties of the fixed-point continuation algorithm and its variants for matrix rank minimization. Heuristics for determining the rank of the matrix when its true rank is not known are also proposed. Some of these algorithms are closely related to greedy algorithms in compressed sensing. Numerical results for these algorithms for solving linearly constrained matrix rank minimization problems are reported. Chapters 4 and 5 considers alternating direction type methods for solving composite convex optimization problems. We present in Chapter 4 alternating linearization algorithms that are based on an alternating direction augmented Lagrangian approach for minimizing the sum of two convex functions. Our basic methods require at most O(1/e) iterations to obtain an e-optimal solution, while our accelerated (i.e., fast) versions require at most O(1/ 3 ) iterations, with little change in the computational effort required at each iteration. For more general problem, i.e., minimizing the sum of K convex functions, we propose multiple-splitting algorithms for solving them. We propose both basic and accelerated algorithms with O(1/e) and O(1/ 3 ) iteration complexity bounds for obtaining an e-optimal solution. To the best of our knowledge, the complexity results presented in these two chapters are the first ones of this type that have been given for splitting and alternating direction type methods. Numerical results on various applications in sparse and low-rank optimization, including compressed sensing, matrix completion, image deblurring, robust principal component analysis, are reported to demonstrate the efficiency of our methods.
Read moreEmploying multiple-kernel support vector machines for counterfeit banknote recognition
Employing multiple-kernel support vector machines for counterfeit banknote recognition
A Special Class of Fractional QCQP and Its Applications on Cognitive Collaborative Beamforming
In this paper, we investigate a special class of the fractional quadratically constrained quadratic problem (QCQP) with more than two quadratic constraints. We propose to effectively solve this special fractional QCQP by one or two convex semidefinite programmings (SDPs). For the two SDPs, one is equivalent to the original fractional QCQP with rank-one relaxation from the Charnes-Cooper transformation and the other, exploiting the optimal value of the former SDP, always has rank-one solution, which is optimal to the former SDP. Theoretical analysis shows that our proposed non-iterative SDP-based algorithm achieves the global optimal solution to the special fractional QCQP. In specific scenarios, our proposed non-iterative SDP-based algorithm has lower computational complexity compared to the second-order cone programming (SOCP)-based and constrained concave convex procedure (CCCP)-based iterative algorithms. We apply the proposed non-iterative SDP-based algorithm on two collaborative beamforming problems in cognitive relay networks, specifically, one is the achievable rate region for two-way non-regenerative cognitive relay networks and the other is the achievable secrecy rate for one-way regenerative cognitive relay networks. Simulation results have shown that our proposed non-iterative SDP-based algorithm achieves the same performance as the SOCP-based iterative algorithm. Our proposed algorithm achieves the better performance than the CCCP-based iterative algorithm.
Read moreMultiple Kernel Based Transfer Learning for the Few-Shot Recognition Task in Smart Home Scene
Multiple Kernel Based Transfer Learning for the Few-Shot Recognition Task in Smart Home Scene
MKBoost: A Framework of Multiple Kernel Boosting
Multiple kernel learning (MKL) is a promising family of machine learning algorithms using multiple kernel functions for various challenging data mining tasks. Conventional MKL methods often formulate the problem as an optimization task of learning the optimal combinations of both kernels and classifiers, which usually results in some forms of challenging optimization tasks that are often difficult to be solved. Different from the existing MKL methods, in this paper, we investigate a boosting framework of MKL for classification tasks, i.e., we adopt boosting to solve a variant of MKL problem, which avoids solving the complicated optimization tasks. Specifically, we present a novel framework of Multiple kernel boosting (MKBoost), which applies the idea of boosting techniques to learn kernel-based classifiers with multiple kernels for classification problems. Based on the proposed framework, we propose several variants of MKBoost algorithms and extensively examine their empirical performance on a number of benchmark data sets in comparisons to various state-of-the-art MKL algorithms on classification tasks. Experimental results show that the proposed method is more effective and efficient than the existing MKL techniques.
Read moreMultiple kernel learning using nonlinear lasso
Multiple kernel learning (MKL) is a principled way for kernel fusion for various learning tasks such as classification, clustering, and dimensionality reduction. The least absolute shrinkage and selection operator (Lasso) allows computationally efficient feature selection based on the linear dependence between input features and output values. In this paper, we develop a novel MKL model based on a nonlinear Lasso, that is, the Hilbert–Schmidt independence criterion (HSIC) Lasso. In the proposed model, we first propose the HSIC Lasso‐based MKL formulation, which has a clear statistical interpretation that minimum redundant kernels with maximum dependence on output labels are found and combined, and also that the global optimal solution can be computed efficiently by solving a Lasso optimization problem. After the optimal kernel is obtained, the support vector machine (SVM) is used to select the prediction hypothesis. It is evident that the proposed MKL is a two‐stage kernel learning approach. Extensive experiments on real‐world datasets from the UCI benchmark repository validate the superiority of the proposed model in terms of prediction accuracy. © 2018 Institute of Electrical Engineers of Japan. Published by John Wiley & Sons, Inc.
Read moreMultimodal continuous affect recognition based on LSTM and multiple kernel learning
In this paper, we propose a Long Short-Term Memory Recurrent Neural Network (LSTM-RNN) and multiple kernel learning (MKL) based multi-modal affect recognition scheme (LSTM-MKL). It takes the LSTM-RNN advantage to model the long range dependencies between successive observations, and uses the MKL power to model the non-linear correlations between the inputs and outputs. For each of the affect dimensions (arousal, valence, expectancy, and power), two LSTM-RNN models are trained, one for each modality. In the recognition phase, the audio and visual features are input to the corresponding learned LSTM models, which in turn produce initial estimates of the affect dimensions. The LSTM outputs are further input into a multi-kernel support vector regression (MK-SVR) for the final recognition. Experimental results carried out on the AVEC2012 database, show that compared to the traditional SVR-LLR (Support Vector Machine - local linear regression) or MK-SVR fusion scheme, the proposed LSTM-MKL fusion scheme obtains higher recognition results, with an correlation coefficient (COR) of 0.354, compared to a COR of 0.124 for SVR-LLR, and 0.168 for MK-SVR, respectively.
Read moreEfficient Multiple Kernel Classification Using Feature and Decision Level Fusion
Kernel methods for classification is a well-studied area in which data are implicitly mapped from a lower-dimensional space to a higher dimensional space to improve classification accuracy. However, for most kernel methods, one must still choose a kernel to use for the problem. Since there is, in general, no way of knowing which kernel is the best, multiple kernel learning (MKL) is a technique used to learn the aggregation of a set of valid kernels into a single (ideally) superior kernel. The aggregation can be done using weighted sums of the precomputed kernels, but determining the summation weights is not a trivial task. Furthermore, MKL does not work well with large datasets because of limited storage space and prediction speed. In this paper, we address all three of these multiple kernel challenges. First, we introduce a new linear feature level fusion technique and learning algorithm, GAMKLp. Second, we put forth three new algorithms, DeFIMKL, DeGAMKL, and DeLSMKL, for nonlinear fusion of kernels at the decision level. To address MKL's storage and speed drawbacks, we apply the Nystrom approximation to the kernel matrices. We compare our methods to a successful and state-of-the-art technique called MKL-group lasso (MKLGL), and experiments on several benchmark datasets show that some of our proposed algorithms outperform MKLGL when applied to support vector machine (SVM)-based classification. However, to no surprise, there does not seem to be a global winner but instead different strategies that a user can employ. Experiments with our kernel approximation method show that we can routinely discard most of the training data and at least double prediction speed without sacrificing classification accuracy. These results suggest that MKL-based classification techniques can be applied to big data efficiently, which is confirmed by an experiment using a large dataset.
Read more