The Crimson Bench

Blog / CTO Insights

Engineering Velocity: The Metrics That Actually Matter

Engineering velocity is one of the most discussed and most poorly measured dimensions of software development. The metrics most commonly tracked tell you about activity, not output — and the gap between the two is where productivity is lost.

2025-11-1710 min read

Why Most Engineering Metrics Mislead

The metrics that are easiest to measure in software development — lines of code written, story points completed, number of commits — are inversely correlated with engineering quality in many contexts. Engineers who write more concise, reusable code often produce fewer lines than those who write verbose, repetitive code. Teams that take the time to properly define and estimate work complete more story points than those that rush estimation. Development activity that does not produce customer value is not velocity — it is noise. The measurement problem in engineering is that the outputs that matter — delivered customer value, system reliability, developer capability — are difficult to measure directly. The inputs — time spent, code written, tasks completed — are easy to measure but weakly correlated with outcomes. Organizations that optimize for easy-to-measure inputs while ignoring hard-to-measure outcomes reliably generate the wrong behaviors: teams that maximize their velocity metrics while delivering diminishing customer value. The solution is not to stop measuring engineering — it is to measure what matters. Four State of DevOps metrics, originally published by DORA (DevOps Research and Assessment), represent the current consensus on the most informative engineering velocity metrics: deployment frequency, lead time for changes, change failure rate, and mean time to recovery. These metrics are outcome-oriented, directly connected to business impact, and difficult to game in ways that decouple metric performance from actual performance.

The DORA Metrics: A Deeper Look

Deployment frequency measures how often code is deployed to production. High-performing engineering organizations deploy multiple times per day; low-performing organizations deploy monthly or less frequently. Deployment frequency is a proxy for the maturity of an organization's development practices: organizations that deploy frequently have the automated testing, continuous integration, and deployment automation infrastructure that makes frequent, safe deployment possible. Lead time for changes measures the time from a code change being committed to that change being in production. Lead time reflects the efficiency of the entire development process: how well requirements are defined, how quickly code is reviewed, how robust automated testing is, and how streamlined the deployment process is. Organizations with long lead times have friction distributed across multiple stages of the development process — each stage adding delay that reduces the speed at which the team can respond to business needs. Change failure rate — the percentage of deployments that result in a degraded service or require remediation — measures quality. A high change failure rate means the team is deploying frequently but breaking things frequently, which produces a net negative on reliability. Mean time to recovery measures how quickly the team can restore service when something does break. Together, change failure rate and mean time to recovery define an organization's reliability posture: how often it creates problems and how quickly it resolves them.

Flow Metrics: Measuring the Development Process

DORA metrics measure the outcome of the engineering process. Flow metrics, popularized by Mik Kersten's work on Project to Product, measure the process itself — providing earlier signals about where the development system is working and where it is constrained. The four core flow metrics are flow velocity (the number of features, defects, risks, and debts completed per period), flow efficiency (the percentage of cycle time that work is actively being progressed), flow time (the end-to-end time from work being identified to being delivered), and flow load (the amount of work in progress at any given time). Flow efficiency is the most consistently revealing of these metrics. For most software development organizations, work is actively progressed only 10 to 20 percent of the time it is in the development system — the remaining 80 to 90 percent is waiting: waiting for requirements clarity, waiting for code review, waiting for QA, waiting for deployment approval. The variability in flow efficiency across teams and across work types reveals where the development system has bottlenecks that are limiting velocity. Flow load — work in progress — is inversely related to flow velocity in ways that are counterintuitive to business stakeholders who want more things done simultaneously. Little's Law demonstrates that as work in progress increases, cycle time increases proportionally, holding throughput constant. Organizations that limit work in progress — that focus on finishing work rather than starting it — consistently achieve higher throughput than those with unlimited work in progress.

Engineering Investment Distribution

Engineering organizations allocate their capacity across four types of work: new features, defect remediation, technical debt retirement, and risk reduction (security, compliance, reliability improvements). The distribution of this allocation is one of the most revealing indicators of an engineering organization's health and strategic positioning. High-performing engineering organizations typically allocate 50 to 60 percent of capacity to new features, 20 to 30 percent to technical debt and infrastructure improvement, and 10 to 20 percent to defect remediation and risk reduction. Organizations that allocate 80 percent or more to new features with minimal investment in technical debt or infrastructure are trading short-term feature velocity for long-term delivery capacity — a trade that consistently produces declining velocity over time as debt accumulates and reliability deteriorates. Tracking investment distribution over time reveals whether engineering capacity allocation is shifting in directions that are sustainable. A team that gradually increases the proportion of time spent on defect remediation is accumulating technical debt faster than it is retiring it — a leading indicator of future velocity decline. A team that is increasing the proportion of time spent on infrastructure improvement is investing in future velocity — a leading indicator of improvement to come.

Frequently Asked Questions

Should engineering teams be measured on story points or velocity?

Story points and velocity are useful internal planning tools but poor cross-team measurement metrics. Different teams estimate story points differently, making comparisons meaningless. Using velocity as a management metric creates incentives to inflate estimates. Outcome metrics — deployment frequency, lead time, customer impact — are far more appropriate for management-level evaluation of engineering performance.

How do we improve deployment frequency without sacrificing quality?

Improving deployment frequency without sacrificing quality requires investment in automated testing coverage, continuous integration infrastructure, and deployment automation. Feature flags — which allow code to be deployed to production but activated only for specific users — are particularly powerful: they decouple deployment from release, enabling very frequent deployments while maintaining control over which users see new capabilities.

What is a healthy change failure rate?

According to DORA research, elite engineering organizations have change failure rates below 5 percent. High performers are in the 5 to 10 percent range. Anything above 15 percent indicates systematic quality problems in the development process. The most common root causes of high change failure rates are insufficient automated testing coverage, inadequate review of infrastructure changes, and insufficient staging environment fidelity.

How should engineering metrics be used in performance reviews?

Individual performance metrics should focus on observable behaviors and contributions — code review quality, documentation, collaboration, technical growth — rather than output metrics like deployment count or story points. Individual output metrics create local optimization at the expense of system performance and incentivize gaming. Team-level DORA and flow metrics are appropriate inputs to team retrospectives and planning processes, not to individual performance assessments.

The Crimson Bench · Est. 2002 · Founded in New York City

Deploy an Executive in 48 Hours

Verified corporate accounts only. Ivy League-educated. Flat-rate pricing. 14-day no-cause cancellation.

25,000+ Ivy League Executives · 150,000+ Global Consultants · 48-Hour Deployment