SRE
Site Reliability Engineering—a discipline that applies software engineering practices to operations problems, managing production systems through code, automation, and data-driven reliability targets.
Full Definition
Site Reliability Engineering (SRE) is a discipline pioneered at Google that applies software engineering principles to operations and infrastructure management. Where traditional operations teams manage systems through manual procedures, SRE teams automate everything that can be automated, define reliability requirements as code (Service Level Objectives), and use data-driven error budgets to balance feature development velocity with reliability investment. The SRE approach eliminates the traditional development-operations conflict by giving the same team (SREs) accountability for both: SREs build the reliability infrastructure that enables developers to ship safely, and they enforce reliability standards that prevent feature development from eroding system stability. The foundational SRE concepts are SLOs (Service Level Objectives—the target reliability level for each service, defined as a measurable metric like 99.9% availability or 100ms P99 latency), SLIs (Service Level Indicators—the actual measured values for each reliability dimension), and error budgets (the allowed amount of unreliability within the SLO definition—a 99.9% availability SLO has a 0.1% error budget, approximately 44 minutes per month). The error budget is the governance mechanism that makes SRE practically effective: when the error budget is being consumed faster than the SLO allows, feature development velocity must slow to allow reliability work to restore the budget. When the error budget is healthy, teams can deploy more aggressively. This data-driven balance prevents the extremes of velocity-at-all-costs (ignoring reliability) and stability-at-all-costs (blocking feature delivery). SRE adoption has grown substantially beyond Google, with Microsoft, LinkedIn, Netflix, Amazon, and thousands of other technology companies implementing SRE practices in various forms. The SRE model works best when: services have sufficient user traffic to make statistical reliability measurement meaningful (low-traffic services may not have enough requests to detect reliability issues until they are severe), engineering teams have the capability to implement automation and reliability infrastructure (SRE is an engineering function, not a traditional operations function), and executive leadership is willing to enforce error budget governance (blocking feature releases when reliability is insufficient is culturally difficult without senior leadership commitment).
FAQs
What is the difference between an SRE and a DevOps engineer?
SRE and DevOps are closely related but not identical. DevOps is a cultural philosophy and set of practices for integrating development and operations; SRE is a specific implementation of DevOps principles using engineering rigor. An SRE engineer specifically focuses on system reliability through software engineering: writing automation, defining and measuring SLOs, implementing observability, and managing capacity. A DevOps engineer may have broader responsibilities including CI/CD pipeline management, tooling, and developer experience that overlap with both SRE and platform engineering functions.
How do you establish SLOs for a service that hasn't had formal reliability requirements before?
Start with user research: what reliability level do users actually need? (Not necessarily the highest technically achievable—99.999% SLO has very different cost implications than 99.9%.) Review historical availability data to understand what the service has actually delivered. Define SLIs first (what metrics indicate the user's experience—request success rate, latency percentiles, error rate?) then set SLO targets as a realistic improvement over current performance rather than an idealized target. Most teams start with a conservative SLO and tighten it as reliability infrastructure matures and they better understand what drives user-impacting failures.
Relevant Executive Roles
The Crimson Bench · Est. 2002 · Founded in New York City
Deploy an Executive in 48 Hours
Verified corporate accounts only. Ivy League-educated. Flat-rate pricing. 14-day no-cause cancellation.
25,000+ Ivy League Executives · 150,000+ Global Consultants · 48-Hour Deployment