Discovering transcriptional modules by Bayesian data integration

Motivation: We present a method for directly inferring transcriptional modules (TMs) by integrating gene expression and transcription factor binding (ChIP-chip) data. Our model extends a hierarchical Dirichlet process mixture model to allow data fusion on a gene-by-gene basis. This encodes the intui...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Bioinformatics 2010-06, Vol.26 (12), p.i158-i167
Hauptverfasser:	Savage, Richard S., Ghahramani, Zoubin, Griffin, Jim E., de la Cruz, Bernard J., Wild, David L.
Format:	Artikel
Sprache:	eng
Schlagworte:	Bayes Theorem Binding Binding Sites Bioinformatics Data integration Dirichlet problem Gene expression Gene Expression Profiling - methods Genes Mathematical models Modules Multigene Family Oligonucleotide Array Sequence Analysis Saccharomyces cerevisiae Proteins - genetics Saccharomyces cerevisiae Proteins - metabolism Transcription Factors - metabolism
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Motivation: We present a method for directly inferring transcriptional modules (TMs) by integrating gene expression and transcription factor binding (ChIP-chip) data. Our model extends a hierarchical Dirichlet process mixture model to allow data fusion on a gene-by-gene basis. This encodes the intuition that co-expression and co-regulation are not necessarily equivalent and hence we do not expect all genes to group similarly in both datasets. In particular, it allows us to identify the subset of genes that share the same structure of transcriptional modules in both datasets. Results: We find that by working on a gene-by-gene basis, our model is able to extract clusters with greater functional coherence than existing methods. By combining gene expression and transcription factor binding (ChIP-chip) data in this way, we are better able to determine the groups of genes that are most likely to represent underlying TMs. Availability: If interested in the code for the work presented in this article, please contact the authors. Contact: d.l.wild@warwick.ac.uk Supplementary information: Supplementary data are available at Bioinformatics online.
ISSN:	1367-4803 1460-2059 1367-4811
DOI:	10.1093/bioinformatics/btq210