Recognizing Shared Operational Events Across Data Center Racks from Distributed Telemetry
编号:87
访问权限:仅限参会人
更新:2026-10-04 23:37:08 浏览:13次
Online
摘要
Modern data centers generate large volumes of component level telemetry, while operational decisions are made at the level of higher-level events. This paper presents a telemetry driven framework for recognizing shared operational events across distinct physical compute racks using publicly available production telemetry from the Marconi100 Tier-0 HPC system. Normal-to-non-normal node-state transitions are aggregated within racks, correlated across racks, and consolidated into operational episodes. Cross-rack synchronization is statistically evaluated using an exact rack-preserving whole-week alignment test, while telemetry trajectories characterize event behavior around onset. Across 45 nodes in three racks, 5,024 valid anomaly onsets produced 77 synchronized cross-rack timestamps and 20 operational episodes. Among the 15,375 alternative relative whole-week rack alignments, the maximum synchronization count was two, compared with 77 in the observed alignment ($p_{\mathrm{exact}}=1/15376\approx6.5\times10^{-5}$), and the episode count remained stable under alternative definitions. Trajectory analysis identified state-dominant or mixed, controlled-transition-like, and abrupt-loss-like modes. Post hoc comparison with CINECA records associated controlled-transition-like events with scheduled operational activity and one abrupt-loss-like event with a documented power-related disruption. Overall, the proposed framework aggregated distributed node-level monitoring anomalies into robust and interpretable shared operational events, providing a higher-level representation for infrastructure monitoring and Artificial Intelligence for IT Operations (AIOps).
关键词
AIOps, anomaly correlation, complex event recognition, data centers, distributed telemetry.
稿件作者
Ilyas Saleem
Amazon Web Services
发表评论