A Statistical Failure/Load Relationship: Results of a Multicomputer Study
- 1 July 1982
- journal article
- Published by Institute of Electrical and Electronics Engineers (IEEE) in IEEE Transactions on Computers
- Vol. C-31 (7) , 697-706
- https://doi.org/10.1109/tc.1982.1676070
Abstract
In this correspondence we present a statistical model which relates mean computer failure rates to level of system activity. Our analysis reveals a strong statistical dependency of both hardware and software component failure rates on several common measures of utilization (specifically CPU utilization, I/O initiation, paging, and job-step initiation rates). We establish that this effect is not dominated by a specific component type, but exists across the board in the two systems studied. Our data covers three years of normal operation (including significant upgrades and reconfigurations) for two large Stanford University computer complexes. The complexes, which are composed of IBM mainframe equipment of differing models and vintage, run similar operating systems and provide the same interface and capability to their users. The empirical data comes from identically structured and maintained failure logs at the two sites along with IBM OS/VS2 operating system performance/load records.Keywords
This publication has 4 references indexed in Scilit:
- WORKLOAD, PERFORMANCE, AND RELlABlLlTY OF DIGITAL COMPUTlNG SYSTEMSPublished by Institute of Electrical and Electronics Engineers (IEEE) ,2005
- On Evaluating the Performability of Degradable Computing SystemsIEEE Transactions on Computers, 1980
- The measurement and management of software reliabilityProceedings of the IEEE, 1980
- The Error Latency of a Fault in a Sequential Digital CircuitIEEE Transactions on Computers, 1976