GrayWulf: Scalable Software Architecture for Data Intensive Computing

Big data presents new challenges to both cluster infrastructure software and parallel application design. We present a set of software services and design principles for data intensive computing with petabyte data sets, named GrayWulf. These services are intended for deployment on a cluster of commo...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Hauptverfasser:	Simmhan, Y., Barga, R., van Ingen, C., Nieto-Santisteban, M., Dobos, L., Li, N., Shipway, M., Szalay, A.S., Werner, S., Heasley, J.
Format:	Tagungsbericht
Sprache:	eng
Schlagworte:	Software architecture
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Big data presents new challenges to both cluster infrastructure software and parallel application design. We present a set of software services and design principles for data intensive computing with petabyte data sets, named GrayWulf. These services are intended for deployment on a cluster of commodity servers similar to the well-known Beowulf clusters. We use the Pan-STARRS system currently under development as an example of the architecture and principles in action.
ISSN:	1530-1605 2572-6862
DOI:	10.1109/HICSS.2009.235