Before you measure AI inputs, prove the outputs

Will 2026 be the year when established companies understand that token consumption as a KPI isn’t good?

AI has been improving considerably over the past year. The hype continues to be deafening and the board starts to push harder for faster AI adoption.

Whilst dealing with the daily operations you come up with a perfect plan: Let’s measure adoption by counting tokens consumption.

Bad idea.

Just like it was a bad idea to measure lines of code (LOC) as an indicator of developer productivity.

It’s such a well-known bad idea that it was first introduced back in the 70s. Goodhart’s Law to be specific.

«When a measure becomes a target, it ceases to be a good measure»

Amazon built «Kirorank», and Meta «Claudeonomics». Both shut down this year after realizing that people would spend tokens on irrelevant problems, simply to satisfy the game.

But regardless of how quotable and clear Mr. Goodhart may be, we all fall into this same trap at least once or twice.

And I see AI adoption as a perfect scenario to fall into the trap of measuring input rather than output.

This article aims to showcase how I see adoption measurements, and when is good to track inputs.

In a nutshell, I believe you should only track inputs AFTER asserting that the foundations are set. Before that, you’re better off tracking outputs.

If Tim used to take 3 weeks to push a new functionality to prod, how much time does it take now that they are equipped with AI?

That is an objective answer related to output, that will give you significantly more understanding of how AI impacts your team than measuring tokens.

Say your team used to have a velocity of 30 points per sprint, and post AI adoption is delivering 20… that’s your cue that something’s off.

The very same output metrics that you used to track before AI are the ones you should use now.

Because it gives you a direct comparison point.

Only after you confirm that AI has proven useful to your people you may start measuring inputs.

Because measuring inputs is optimizing.

You shouldn’t optimize that which you don’t even know whether is worth it in the first place.

AI is a multiplier, of both good and bad. Google’s DORA has many reports about this if you’re interested.

This means that ensuring FIRST that your team is benefiting from having AI around is far more important than assuming that everyone needs to hit a certain amount of tokens.

This should give you enough to think about regarding the company-wide strategy. But what about individuals?

I run our 1:1 meetings, and something I keep asking myself is how do I keep individuals accountable without pointing to a number.

The answer is the same. Measure the observable outputs.

«It’s been 3 weeks since you got your AI license and your time to prod stayed the same, your regression cycle with QA is the same, and your time doing peer reviews is the same. What impact has AI had in this?»

And you take it from there. Ask without assumptions that AI will magically fit everyone, and let them tell you how they been using it.

If the answer is that they haven’t been using it, then that’s your cue to push harder.

If they’ve been using it, then you may need to dig deeper to understand if it’s a process issue or simply a bad batch of tasks to test on.

Whatever the case, know that AI isn’t magical and won’t make all your people 10x better overnight.

And know that, as you start, measuring observable outputs is better than forcing inputs.


Originally posted at: https://www.linkedin.com/pulse/before-you-measure-ai-inputs-prove-outputs-mauricio-guzman-muletaber-5busf

Otras

publicaciones

Antes de medir los insumos de la IA, demostrá sus resultados

¿Será 2026 el año en que las empresas establecidas entiendan que el consumo de tokens como

Más allá de la ruta más corta: el desafío de optimizar la última milla

Esta nota adapta aprendizajes operativos publicados originalmente por Francisco Crizul, CTO de Deri, nuestra empresa de

Más allá del picking: cómo optimizar una ola en un depósito real (y en supermercados)

Esta nota adapta aprendizajes operativos publicados originalmente por Francisco Crizul, CTO de Deri, nuestra empresa de