SLO, SLA, and SLI
Three interconnected reliability concepts: SLIs are metrics measuring system behavior, SLOs are internal reliability targets, and SLAs are external contractual commitments with consequences for failure.
Full Definition
SLI (Service Level Indicator), SLO (Service Level Objective), and SLA (Service Level Agreement) form a layered reliability management framework. An SLI is the raw metric measuring a specific dimension of service behavior—request success rate (ratio of successful requests to total requests), latency (time to serve a request), throughput (requests per second), and error rate (ratio of error responses to total responses) are common SLIs. SLIs are the factual measurements; they don't express a judgment about what level of performance is acceptable. An SLO is the internal target for an SLI—a performance commitment the team makes to itself about what reliability level they aim to maintain. A team might set an SLO of "99.9% of requests succeed in less than 200ms" for their API service. SLOs define the reliability standard that drives engineering and operational decisions: when SLO performance is healthy, teams can invest in feature development; when SLO performance is at risk, reliability work takes priority over new features. SLOs should reflect what users actually need—not the highest achievable level, but the level below which users experience meaningful degradation. Well-designed SLOs have error budgets (the allowed non-compliance within the SLO target) that governance teams use to make deployment and investment decisions. An SLA is the external contractual commitment made to customers—typically more conservative than internal SLOs to provide buffer between internal targets and contractual obligations, and carrying defined consequences (financial credits, termination rights) when the SLA is not met. A company might maintain a 99.9% availability SLO internally while committing to 99.5% availability in customer contracts, providing a 40-minute/month buffer between internal performance targets and contractual obligations. The SLA is what the customer negotiates and what the legal team drafts; the SLO is what engineering operates to. SLOs that significantly exceed SLAs indicate over-engineering for the committed standard; SLOs that equal SLAs create no buffer for the inevitable gap between targets and contractual commitments.
FAQs
How many SLOs should a service have?
Most services should have 2-5 SLOs covering the dimensions of reliability most important to users: availability (can users access the service at all?), latency (how fast does the service respond?), error rate (how often do requests fail?), and for data services, data freshness (how current is the data?). More than 5-7 SLOs per service creates governance complexity without proportionally improving user experience measurement. The goal is a small set of SLOs that comprehensively capture whether users are having a good experience—not an exhaustive list of all technical system metrics.
What should a company do when it misses its SLA?
When an SLA breach occurs: immediately notify affected customers per the agreement's notification requirements, calculate the credit or remedy owed under the SLA terms and proactively apply it (don't wait for customers to request it), conduct a post-incident review to understand root cause and implement prevention measures, and report the breach and prevention measures to customers. Proactive communication and credit application—before customers ask—preserves relationships despite the reliability failure. Companies that respond to SLA breaches with silence or customer-initiated credit requests generate disproportionate churn relative to the underlying reliability failure.
Relevant Executive Roles
The Crimson Bench · Est. 2002 · Founded in New York City
Deploy an Executive in 48 Hours
Verified corporate accounts only. Ivy League-educated. Flat-rate pricing. 14-day no-cause cancellation.
25,000+ Ivy League Executives · 150,000+ Global Consultants · 48-Hour Deployment