2 min readcostengineering-productivitymetrics

Token counts are about to become the new lines of code

We already learned that measuring commits and lines of code doesn't measure effectiveness. Token-spend dashboards are about to teach us the same lesson again, with a dollar sign attached.

Chris RaethkeCo-founder & CTO, Cadence

We’ve made this mistake before. We measured commits and lines of code, and learned the hard way that activity isn’t effectiveness. Now the first token-spend dashboards are showing up across engineering orgs, and it’s the same mistake again - just with a dollar sign attached this time, which makes it more dangerous, not less.

The gap nobody’s tracking

Benedict Evans published a long piece recently on token pricing. Buried in it is a line that matters more than any of the pricing charts around it: nobody knows how much of the current usage surge has an ROI that can actually be quantified to a CFO.

That’s the actual gap. Not the price of tokens - what the tokens bought.

I see it inside teams every week. One developer burns through an entire Max plan and ships nothing worth mentioning. Another spends more and turns three weeks of work into four days. A token leaderboard says those two are basically the same person, ranked by spend. They are not remotely the same person, and treating them as comparable is worse than not measuring at all - it actively points you at the wrong conclusion.

The uncomfortable part

Once you look at output instead of spend, the obvious next move gets uncomfortable fast: the right move for your best engineers is probably to spend more on tokens, not less. If someone is reliably converting token spend into three weeks of work in four days, throttling their usage to hit a per-seat budget target is a strictly worse outcome for the business than letting them spend freely and throttling someone whose spend isn’t converting into anything.

Most token-governance conversations I’ve seen so far are structured the wrong way round: cap everyone at roughly the same number, because that’s administratively simple and looks fair on a spreadsheet. It isn’t fair, and it isn’t even efficient - it just optimises for a number that’s easy to see instead of the number that actually matters.

Where this goes

Evans’s read is that tokens end up as commodity infrastructure, with the value captured further up the stack - the same shape cloud compute took, and the same shape most raw infrastructure inputs take eventually. If he’s right, price per million tokens becomes a rounding error in most engineering budgets, and the only number that will have mattered the whole time is the one almost nobody tracks today: cost per bug fixed, cost per feature that actually shipped.

If your CFO asked what last month’s AI spend actually produced, could you answer that - not with a token count, but with an outcome?