Measuring productivity isn’t something you invent. It’s something you adopt and adapt

Marcelo Scheidt
Marcelo Scheidt
Verified Author Verified Author
3 September

In the previous article, I argued that AI does not create productivity and only amplifies what a team already is, and that the right question was therefore never how much AI sped us up, but whether the team is delivering what was agreed, on time and with quality. What was missing is the uncomfortable part, which is how to measure that in practice, especially in a consultancy, where every project lives in a different world.

The first temptation is usually to build a metric of your own, from scratch, tailor made. And it is almost always a mistake.

Why adopt instead of invent

The industry has spent more than a decade studying this problem, and the knowledge is already consolidated into a few frameworks worth knowing before reinventing the wheel. DORA was born to measure delivery speed and stability and is the most objective of them all. SPACE came to remind us that productivity is multidimensional and that satisfaction and collaboration also count, even if it leaves the job of defining the metrics to you. DevEx brought the focus to the experience of the person developing, to cognitive load and the state of flow. And DX Core 4 pulled the three together into a practical model organized around four dimensions that balance each other, namely speed, effectiveness, quality, and impact.

Adopting one of these is not laziness, it is the opposite. It means starting from a point validated by thousands of teams instead of discovering on your own, through error, what is already known. The real work is not in creating the framework, it is in adapting it to your context, which is exactly where most companies slip.

A consultancy’s context changes everything

These frameworks were designed with product companies in mind, with stable teams delivering continuously to their own users. A consultancy does not work that way, and ignoring that difference is the shortest path to a pretty dashboard that says nothing.

Here the teams are ephemeral, assembled per project, with different stacks and maturity levels for each client, so the unit worth tracking stops being the team and becomes the project. And the business impact dimension that closes Core 4 is simply not yours, because the production deploy, the business indicator, and the revenue belong to the client. What remains under your control is what actually matters commercially, and what I would call contractual impact, meaning delivering the agreed scope within the time that was sold, without the work coming back as rework, and being able to repeat that project after project. Predictability and quality, not volume.

The data belongs to the client, and that is a real problem

There is still a constraint that almost no one anticipates and that decides the choice of tool before any comparison of price. In a consultancy, much of the work happens in the client’s Git, with the tools the client operates, which means the engineering data you would like to measure often is not even yours.

And that is where the care lies. Pointing an analytics tool at a client’s private repository without explicit contractual consent is a legal and trust risk, not a technical detail, and sending that data to a third party service runs into NDAs, security policy, and, depending on the sector, compliance. That is why, in our case, a self hosted and open source solution like Apache DevLake tends to be more suitable than the commercial platforms, not because of price, but because the data can stay in a controlled environment, without leaving, while connecting to the variety of sources that the reality of multiple clients demands.

The good news is that the backbone of the measurement does not depend on the client. Predictability is measured with what is already yours, such as estimated against actual, adherence to deadlines, planned scope against delivered scope, and margin, and all of that lives in your project management and in the contracts. The signal coming from the client’s Git enriches the reading, but it comes later, project by project, wherever there is access and permission.

And where does AI fit

Here I return to the point from the previous article, now with a practical consequence. If perception deceives, then no serious measurement of AI impact can rest on asking people whether AI helped, because that measures adoption and satisfaction, never real gain.

To measure impact for real, three cautions help. The first is to separate by type of task, since AI performs very well on repetitive code, migrations, and exploring new territory, and performs poorly on complex work over a mature base, so that the average between the two ends up hiding both. The second is to compare cohorts instead of believing in easy causality, observing projects with high and low AI adoption across the same delivery and stability metrics. The third is to keep stability at the center of the reading, because the pattern that repeats across the industry is individual output rising while delivery and quality fail to keep up, and in a consultancy that has a name, which is eroded margin and an unhappy client.

Where to start

None of this needs to become a giant program all at once, and the path that works is one of maturity, not of a big bang. Start by deciding, together with leadership, which outcome you want to drive, because measuring without that agreement produces dashboards no one uses. Instrument first the data that is already yours and establish a baseline before attributing anything to AI, since without a baseline there is no comparison, only guesswork. Then add the human layer, a light and periodic pulse on the team’s experience, which is what turns a number into a conversation and protects against the game of chasing a metric. Only then overlay the AI reading, with adoption first, impact next, and cost last.

And above all, make it clear to everyone, in writing, that a metric serves to steer, not to score. The moment a number becomes an individual performance grade, it stops describing reality and starts being managed. In a consultancy, where trust is the asset, that care is not optional.

In the end, measuring productivity with AI is more a matter of discipline than of tooling. Technology collects the data, but it is clarity about what matters that turns data into decisions. Those who start with the tool end up with numbers. Those who start with the right question end up with answers.

Published on: Article
Marcelo Scheidt
Marcelo Scheidt
Verified AuthorVerified Author

Senior Principal Engineer at Zallpy, a software engineer with over 15 years of experience in systems engineering, architecture, and infrastructure, working on defining technical vision and architectural strategies for complex, large-scale solutions. He currently serves as Senior Principal Engineer at Zallpy, where he leads the architectural evolution across multiple domains, supports executives in technical decision-making, conducts maturity and risk assessments, and develops technical leadership by mentoring senior engineers, consistently connecting cloud-native architectures, modern DevOps practices, and engineering excellence to business objectives.

Senior Principal Engineer at Zallpy, a software engineer with over 15 years of experience in systems engineering, architecture, and infrastructure, working on defining technical vision and architectural strategies for complex, large-scale solutions. He currently serves as Senior Principal Engineer at Zallpy, where he leads the architectural evolution across multiple domains, supports executives in technical decision-making, conducts maturity and risk assessments, and develops technical leadership by mentoring senior engineers, consistently connecting cloud-native architectures, modern DevOps practices, and engineering excellence to business objectives.