Bayesian Model-Based Clustering Procedures

This article establishes a general formulation for Bayesian model-based clustering, in which subset labels are exchangeable, and items are also exchangeable, possibly up to covariate effects. The notational framework is rich enough to encompass a variety of existing procedures, including some recent...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Journal of computational and graphical statistics 2007-09, Vol.16 (3), p.526-558
Hauptverfasser:	Lau, John W, Green, Peter J
Format:	Artikel
Sprache:	eng
Schlagworte:	Algorithms Bayesian analysis Bayesian networks Datasets Dirichlet process Galaxy clusters Heuristic Hierarchical clustering Integer programming Leukemia Loss functions Mathematical procedures Modeling Multilevel models Objective functions Optimization algorithms Point estimators Stochastic models Stochastic search Studies
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	This article establishes a general formulation for Bayesian model-based clustering, in which subset labels are exchangeable, and items are also exchangeable, possibly up to covariate effects. The notational framework is rich enough to encompass a variety of existing procedures, including some recently discussed methods involving stochastic search or hierarchical clustering, but more importantly allows the formulation of clustering procedures that are optimal with respect to a specified loss function. Our focus is on loss functions based on pairwise coincidences, that is, whether pairs of items are clustered into the same subset or not. Optimization of the posterior expected loss function can be formulated as a binary integer programming problem, which can be readily solved by standard software when clustering a modest number of items, but quickly becomes impractical as problem scale increases. To combat this, a new heuristic item-swapping algorithm is introduced. This performs well in our numerical experiments, on both simulated and real data examples. The article includes a comparison of the statistical performance of the (approximate) optimal clustering with earlier methods that are model-based but ad hoc in their detailed definition.
ISSN:	1061-8600 1537-2715
DOI:	10.1198/106186007X238855