Tracking Probabilistic Correlation of Monitoring Data for Fault Detection in Complex Systems
- 21 July 2006
- conference paper
- Published by Institute of Electrical and Electronics Engineers (IEEE)
- Vol. 1 (15300889) , 259-268
- https://doi.org/10.1109/dsn.2006.70
Abstract
Due to their growing complexity, it becomes extremely difficult to detect and isolate faults in complex systems. While large amount of monitoring data can be collected from such systems for fault analysis, one challenge is how to correlate the data effectively across distributed systems and observation time. Much of the internal monitoring data reacts to the volume of user requests accordingly when user requests flow through distributed systems. In this paper, we use Gaussian mixture models to characterize probabilistic correlation between flow-intensities measured at multiple points. A novel algorithm derived from expectation-maximization (EM) algorithm is proposed to learn the "likely" boundary of normal data relationship, which is further used as an oracle in anomaly detection. Our recursive algorithm can adaptively estimate the boundary of dynamic data relationship and detect faults in real time. Our approach is tested in a real system with injected faults and the results demonstrate its feasibilityKeywords
This publication has 7 references indexed in Scilit:
- Discovering likely invariants of distributed transaction systems for autonomic system managementCluster Computing, 2006
- Recursive unsupervised learning of finite mixture modelsPublished by Institute of Electrical and Electronics Engineers (IEEE) ,2004
- Diagnosis of asynchronous discrete-event systems: a net unfolding approachIEEE Transactions on Automatic Control, 2003
- An Automated Fault Diagnosis System Using Hierarchical Reasoning and Alarm CorrelationJournal of Network and Systems Management, 2001
- On-line unsupervised outlier detection using finite mixtures with discounting learning algorithmsPublished by Association for Computing Machinery (ACM) ,2000
- High speed and robust event correlationIEEE Communications Magazine, 1996
- Bayesian Data AnalysisPublished by Taylor & Francis ,1995