Measuring engineering productivity is harder than ever

· Source: LeadDev · Field: Technology & Digital — Software Development & Engineering, Artificial Intelligence & Machine Learning, Data Science & Analytics · Depth: Intermediate, medium

Summary

AI has significantly complicated the measurement of engineering productivity by rendering traditional proxies, such as pull requests and commit counts, less reliable indicators of value. While OpenAI engineers using Codex open 70% more pull requests, this activity increase does not automatically equate to higher business value, as engineers increasingly focus on reviewing AI-generated code, defining architectural intent, and improving underlying systems. The article highlights that metrics like lines of code, commit counts, and story points, which rewarded volume, are now insufficient. Instead, frameworks like DORA and SPACE offer a more holistic view, emphasizing delivery system performance and multidimensional productivity. The shift means value now stems from enhancing the software production "factory" rather than individual code output, leading to a decoupling of visible activity and actual business value. This necessitates evaluating the engineering system's overall effectiveness in delivering customer value, rather than relying on individual output metrics.

Key takeaway

For Directors of AI/ML evaluating engineering team performance, your traditional activity metrics like pull requests are now misleading. You should immediately begin pairing activity metrics with outcome-focused measures, such as change failure rate or customer adoption, to accurately reflect value. Prioritize retiring one activity metric each quarter, replacing it with a credible outcome metric. Additionally, ask senior engineers where their most valuable work, like architecture simplification or incident prevention, isn't captured by Git history. This approach ensures your QBRs reflect true business impact, not just visible code output.

Key insights

AI adoption decouples engineering activity from value, necessitating new metrics focused on systemic improvements over individual code output.

Principles

In practice

Topics

Best for: CTO, Director of AI/ML, VP of Engineering/Data, MLOps Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by LeadDev.