Adaptive sparseness for supervised learning

Top Cited Papers

8 September 2003

journal article
Published by Institute of Electrical and Electronics Engineers (IEEE)

Vol. 25 (9) , 1150-1159
https://doi.org/10.1109/tpami.2003.1227989

Abstract

The goal of supervised learning is to infer a functional mapping based on a set of training examples. To achieve good generalization, it is necessary to control the "complexity" of the learned function. In Bayesian approaches, this is done by adopting a prior for the parameters of the function being learned. We propose a Bayesian approach to supervised learning, which leads to sparse solutions; that is, in which irrelevant parameters are automatically set exactly to zero. Other ways to obtain sparse classifiers (such as Laplacian priors, support vector machines) involve (hyper)parameters which control the degree of sparseness of the resulting classifiers; these parameters have to be somehow adjusted/estimated from the training data. In contrast, our approach does not involve any (hyper)parameters to be adjusted or estimated. This is achieved by a hierarchical-Bayes interpretation of the Laplacian prior, which is then modified by the adoption of a Jeffreys' noninformative hyperprior. Implementation is carried out by an expectation-maximization (EM) algorithm. Experiments with several benchmark data sets show that the proposed approach yields state-of-the-art performance. In particular, our method outperforms SVMs and performs competitively with the best alternative techniques, although it involves no tuning or adjustment of sparseness-controlling hyperparameters.

Keywords

This publication has 25 references indexed in Scilit:

Wavelet-based image estimation: an empirical Bayes approach using Jeffrey's noninformative prior
IEEE Transactions on Image Processing, 2001
A new approach to variable selection in least squares problems
IMA Journal of Numerical Analysis, 2000
Analysis of multiresolution image denoising schemes using generalized Gaussian and complexity priors
IEEE Transactions on Information Theory, 1999
Atomic Decomposition by Basis Pursuit
SIAM Journal on Scientific Computing, 1998
Bayesian classification with Gaussian processes
Published by Institute of Electrical and Electronics Engineers (IEEE) ,1998
Bayesian Regularization and Pruning Using a Laplace Prior
Neural Computation, 1995
Normal/Independent Distributions and Their Applications in Robust Regression
Journal of Computational and Graphical Statistics, 1993
Bayesian Analysis of Binary and Polychotomous Response Data
Journal of the American Statistical Association, 1993
Networks for approximation and learning
Proceedings of the IEEE, 1990
Ridge Regression: Biased Estimation for Nonorthogonal Problems
Technometrics, 1970