A Large-Scale Study of Failures on Petascale Supercomputers

With the rapid development of supercomputers, the scale and complexity are ever increasing, and the reliability and resilience are faced with larger challenges. There are many important technologies in fault tolerance, such as proactive failure avoidance technologies based on fault prediction, react...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Journal of computer science and technology 2018, Vol.33 (1), p.24-41
Hauptverfasser:	Liu, Rui-Tao, Chen, Zuo-Ning
Format:	Artikel
Sprache:	eng
Schlagworte:	Analysis Artificial Intelligence Computer engineering Computer Science Data Structures and Information Theory Failure Failure analysis Failure times Fault diagnosis Fault tolerance Faults Information Systems Applications (incl.Internet) R&D Random access memory Regular Paper Reliability Research & development Software Engineering Supercomputers Theory of Computation
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	With the rapid development of supercomputers, the scale and complexity are ever increasing, and the reliability and resilience are faced with larger challenges. There are many important technologies in fault tolerance, such as proactive failure avoidance technologies based on fault prediction, reactive fault tolerance based on checkpoint, and scheduling technologies to improve reliability. Both qualitative and quantitative descriptions on characteristics of system faults are very critical for these technologies. This study analyzes the source of failures on two typical petascale supercomputers called Sunway BlueLight (based on multi-core CPUs) and Sunway TaihuLight (based on heterogeneous manycore CPUs). It uncovers some interesting fault characteristics and finds unknown correlation relationship among main components’ faults. Finally the paper analyzes the failure time of the two supercomputers in various grains of resource and different time spans, and builds a uniform multi-dimensional failure time model for petascale supercomputers.
ISSN:	1000-9000 1860-4749
DOI:	10.1007/s11390-018-1806-7