Measurement
Measuring engineering work when AI writes code
When a tool can produce a thousand lines in a minute, counting output tells you almost nothing. Engineering leads need measures of delivery, quality and team health that AI can't inflate.
Metrics that mislead
- Lines of code or commits. Always weak, now meaningless. More code is often more to maintain.
- AI suggestion acceptance rate. It measures how often people press accept, not whether the result was right. Useful to a tool vendor, not as a team goal.
- Pull requests per engineer. Easy to raise by splitting work or generating trivial changes.
- Story points completed. An estimation aid, not a productivity measure, and easily re-scaled.
Any metric that becomes a target will be optimized. With AI tools, optimizing output metrics has never been easier, which is exactly why they shouldn't be targets.
Delivery: the DORA measures
The DORA research program's software delivery measures remain a sound starting point because they reward shipping safely, not shipping a lot. DORA now lists five, grouped as throughput and instability:
| Measure | What it tells you |
|---|---|
| Change lead time | How long a change takes from commit to running in production |
| Deployment frequency | How often the team deploys change |
| Failed deployment recovery time | How quickly the team recovers when a deployment fails |
| Change fail rate | How often a deployment needs immediate intervention |
| Deployment rework rate | How often deployments are unplanned fixes for production incidents |
If AI tools raise deployment frequency while change fail rate or rework rate climbs, the team is going faster, not better. Watch them together. DORA's definitions, and how they compare with SPACE, DevEx and DX Core 4, are in DORA vs SPACE vs DX Core 4.
People and flow: the SPACE framework
The SPACE framework argues that developer productivity can't be captured in one number. It looks across Satisfaction and well-being, Performance, Activity, Communication and collaboration, and Efficiency and flow. Use a few measures from different dimensions, and include what engineers themselves report. Activity is the dimension AI tools inflate most easily, so never read it alone.
Measures worth adding in the AI era
- Review load and review time. If change volume rises faster than review capacity, quality will follow it down.
- Rework and reverts. Changes reverted or substantially rewritten within weeks of merging.
- Escaped defects and incidents, linked back to the changes that caused them.
- Onboarding time to first meaningful change, and whether new engineers can explain it.
- Engineer-reported friction: short, regular surveys on what slows the team down.
Running an honest AI tool evaluation
- Decide what outcome you expect before rollout, and measure a baseline first.
- Compare like with like: similar work, similar teams, over enough time to see quality effects, not just speed.
- Count the costs: licenses, review time, incidents and security work, not just time saved drafting.
- Ask engineers. Their experience of where the tool helps and where it hurts is data.
- Don't rely on perception alone. In a randomized trial published by the research group METR in July 2025, experienced open-source developers took longer on tasks when AI tools were allowed while believing the tools had sped them up. METR said in February 2026 that the slowdown likely no longer applies to current tools, but the gap between perceived and measured speed is the reason to measure both.
A simple team scorecard
| Area | Measure | Source | Review |
|---|---|---|---|
| Delivery | Change lead time, deployment frequency | Pipeline and version control | Monthly |
| Stability | Change fail rate, rework rate, recovery time | Deployment and incident records | Monthly |
| Review health | Time to first review, open pull request age | Code host | Every two weeks |
| Experience | Short survey on friction, focus and tools | Engineers | Quarterly |
| Impact | Share of time on new capabilities vs unplanned work | Work tracker | Quarterly |
Review trends with the team, not just with leadership. The conversation about why a number moved is where most of the value is. When stability measures slip, the causes often show up in postmortems and in the technical debt register.
Use metrics on systems, not individuals. Team-level delivery measures improve conversations about process. Individual output rankings damage trust and are especially easy to game with AI tools.
Common questions
How do you measure the impact of AI coding tools?
Take a baseline of delivery, stability and experience measures before rollout, then compare over several months. Include costs such as review time and incidents, not only drafting time saved.
Is AI suggestion acceptance rate a useful metric?
It shows how often suggestions were accepted, not whether the result was right. It can help a tool evaluation; it should not be a team goal.
Should I measure individual developer productivity?
Team-level measures are more reliable and less damaging. Individual performance is better judged through the behaviors in your career ladder.
What if review becomes the bottleneck?
Limit work in progress, keep changes small and grow more reviewers. See reviewing AI-generated code.
Last reviewed 2026-09-17