EngineeringLead

Measurement

Measuring engineering work when AI writes code

When a tool can produce a thousand lines in a minute, counting output tells you almost nothing. Engineering leads need measures of delivery, quality and team health that AI can't inflate.

Metrics that mislead

Any metric that becomes a target will be optimized. With AI tools, optimizing output metrics has never been easier, which is exactly why they shouldn't be targets.

Delivery: the DORA measures

The DORA research program's software delivery measures remain a sound starting point because they reward shipping safely, not shipping a lot. DORA now lists five, grouped as throughput and instability:

MeasureWhat it tells you
Change lead timeHow long a change takes from commit to running in production
Deployment frequencyHow often the team deploys change
Failed deployment recovery timeHow quickly the team recovers when a deployment fails
Change fail rateHow often a deployment needs immediate intervention
Deployment rework rateHow often deployments are unplanned fixes for production incidents

If AI tools raise deployment frequency while change fail rate or rework rate climbs, the team is going faster, not better. Watch them together. DORA's definitions, and how they compare with SPACE, DevEx and DX Core 4, are in DORA vs SPACE vs DX Core 4.

People and flow: the SPACE framework

The SPACE framework argues that developer productivity can't be captured in one number. It looks across Satisfaction and well-being, Performance, Activity, Communication and collaboration, and Efficiency and flow. Use a few measures from different dimensions, and include what engineers themselves report. Activity is the dimension AI tools inflate most easily, so never read it alone.

Measures worth adding in the AI era

  1. Review load and review time. If change volume rises faster than review capacity, quality will follow it down.
  2. Rework and reverts. Changes reverted or substantially rewritten within weeks of merging.
  3. Escaped defects and incidents, linked back to the changes that caused them.
  4. Onboarding time to first meaningful change, and whether new engineers can explain it.
  5. Engineer-reported friction: short, regular surveys on what slows the team down.

Running an honest AI tool evaluation

A simple team scorecard

AreaMeasureSourceReview
DeliveryChange lead time, deployment frequencyPipeline and version controlMonthly
StabilityChange fail rate, rework rate, recovery timeDeployment and incident recordsMonthly
Review healthTime to first review, open pull request ageCode hostEvery two weeks
ExperienceShort survey on friction, focus and toolsEngineersQuarterly
ImpactShare of time on new capabilities vs unplanned workWork trackerQuarterly

Review trends with the team, not just with leadership. The conversation about why a number moved is where most of the value is. When stability measures slip, the causes often show up in postmortems and in the technical debt register.

Use metrics on systems, not individuals. Team-level delivery measures improve conversations about process. Individual output rankings damage trust and are especially easy to game with AI tools.

Common questions

How do you measure the impact of AI coding tools?

Take a baseline of delivery, stability and experience measures before rollout, then compare over several months. Include costs such as review time and incidents, not only drafting time saved.

Is AI suggestion acceptance rate a useful metric?

It shows how often suggestions were accepted, not whether the result was right. It can help a tool evaluation; it should not be a team goal.

Should I measure individual developer productivity?

Team-level measures are more reliable and less damaging. Individual performance is better judged through the behaviors in your career ladder.

What if review becomes the bottleneck?

Limit work in progress, keep changes small and grow more reviewers. See reviewing AI-generated code.

Last reviewed 2026-09-17