Implementing SLI/SLA/SLO Metrics in anASP.NET 8 Microservices Architecture Using Open Telemetry and Prometheus

Abstract

This article presents the results of a study on implementing a service level metrics system (SLI/SLA/SLO) in a high-load web application on the ASP.NET 8 platform. A practical implementation of a monitoring system using OpenTelemetry Metrics and a Prometheus endpoint is considered. A comparative analysis of system performance and reliability indicators before and after implementing the metrics is conducted. The mechanisms by which metrics influence operational indicators are described in detail: identifying hidden performance issues, transforming the alerting system, prioritizing engineering efforts, managing the balance between development speed and reliability through an error budget, and automating response processes. The study results showed an improvement in the average time to detect incidents by 73%, a reduction in service restoration time by 58%, and an increase in overall system availability from 98.2% to 99.7%. A methodology for determining target SLO values based on business requirements and architectural constraints is presented

 

 

PDF