A Metastate HMM with Application to Gene Structure Identification in Eukaryotes
We introduce a generalized-clique hidden Markov model (HMM) and apply it to gene finding in eukaryotes ( C. elegans ). We demonstrate a HMM structure identification platform that is novel and robustly-performing in a number of ways. The generalized clique HMM begins by enlarging the primitive hidden...
Gespeichert in:
Veröffentlicht in: | EURASIP journal on advances in signal processing 2010-01, Vol.2010 (1), Article 581373 |
---|---|
Hauptverfasser: | , |
Format: | Artikel |
Sprache: | eng |
Schlagworte: | |
Online-Zugang: | Volltext |
Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Zusammenfassung: | We introduce a generalized-clique hidden Markov model (HMM) and apply it to gene finding in eukaryotes (
C. elegans
). We demonstrate a HMM structure identification platform that is novel and robustly-performing in a number of ways. The generalized clique HMM begins by enlarging the primitive hidden states associated with the individual base labels (as exon, intron, or junk) to substrings of primitive hidden states, or footprint states, having a minimal length greater than the footprint state length. The emissions are likewise expanded to higher order in the fundamental joint probability that is the basis of the generalized-clique, or "metastate", HMM. We then consider application to eukaryotic gene finding and show how such a metastate HMM improves the strength of coding/noncoding-transition contributions to gene-structure identification. We will describe situations where the coding/noncoding-transition modeling can effectively recapture the exon and intron heavy tail distribution modeling capability as well as manage the exon-start needle-in-the-haystack problem. In analysis of the
C. elegans
genome we show that the sensitivity and specificity (SN,SP) results for both the individual-state and full-exon predictions are greatly enhanced over the standard HMM when using the generalized-clique HMM. |
---|---|
ISSN: | 1687-6180 1687-6172 1687-6180 |
DOI: | 10.1155/2010/581373 |