Burning Tokens

Burning tokens is a double-edged sword. Powerful weapons can either give a lot of leverage or drag you down. It's easy to think that you are not shooting yourself in the foot with unlimited tokens. The thing is, you won't be shooting yourself in the foot; you will be expending a lot of company money and perhaps with lots of amplified waste. But only if you do it wrong or don't know what you are doing, and that's the very tricky part.

Big Tech and big companies already have unsustainable token usage, like Uber. The list goes on and on: Microsoft, Walmart, Meta, JPMorgan, AT&T, and many others. For some, tokenmaxxing is over. Is it? Really depends on the cost of inference. Jev is showing that a classifier can be powerful and reliable, at least in the sense of low cost and speed, not saying anything about accuracy, which is not the score.

The problem is not using tokens; that is the symptom. The problem is a lack of training; there is a huge AI enablement GAP in the industry. It's very easy to produce more PRs and get overwhelmed; that's what AI sets you up for.

Historically, companies never focused on learning, and at least before AI agents, you were forced to know more of what was going on. Yes, it was still possible to not understand everything and do a poor job. Now it's much easier to do a poor job and understand nothing.

Bigger model => Smaller Model

It's not good. Leads to a waterfall trap. Same anti-patterns and SDD and waterfall. Planning with Fable/Opus and executing with Sonnet is not a good approach. Big companies do it because they have to limit the budget evenly across teams. In this sense, startups have real advantages because they can have unlimited tokens since they are much smaller and have much less software. Big companies have much more software and do not modernize aggressively as they should, so they need to use sub-optimal and even anti-pattern strategies.

Model routing is good and works; look, Jev, but the real problem here is forcing a rigid planning/execution split because of budget constraints, and then they have an AI waterfall.

Burning Tokens the right way

It's possible to get much more done and speed up some queues, not all, and not the whole SDLC. But first, you need training; your team and your organization need to get trained.

Usually, token burning happens in ways you would not use an LLM; these techniques do not remove 100% of the need to review and understand, but they can give you speed:

  • Using tokens to monitor PRs: Comments, Github Actions outputs, output from automation or humans.
  • Manual Work must become automation: If you do a manual code review and discover 5-10 things you would write as comments, you need to turn that into linter rules, unit tests, and automation; otherwise, you are just being a bottleneck.
  • Merging is also a bottleneck: What you need is a scoring system for parts of your system or services. If it's slow risk, allow merge without review; it's high risk and demands review.

Watch out for your mentality

It's easy for humans to repeat the same behaviors over and over. If you are used to doing code reviews, you will keep catching the same things forever, so the non-obvious thing is to extrapolate patterns and move to automation. If people never talk about how they are using AI, something even bigger is broken. People must talk at least weekly about what they are doing, what's working, and what's not.

Because you can type with Claude-Code, it does not mean your team is functional and things are working. Shipping is never a sign of the right culture and solid practices; dysfunctional organizations and companies existed long before AI, and now AI amplifies that too.

The Style Trap

Code review has always had a style trap: arbitrary style decisions like whether this should be an arrow function or not, whether to use a stream or a for loop, etc. It's easy for engineers to be sucked into "coding" and miss patterns, missing architecture and design, especially if the culture is just ship and never talk, never review, never justify why. Why we do things matters, always matters, and still matters.

Soul Sucking

The anti-pattern that burning tokens can lead to is removing all communication, all retrospection, all design, all architecture, and just shipping. Deliberate or not, the result is technical debt at scale, called Slop today. It's lack of understanding, it's lack of maintainability, and perhaps the worst from us as humans. Some call this soul-sucking, and it's real and bad.

The goal is not to burn fewer tokens either. The goal is to spend machine intelligence on removing scarce human bottlenecks without automating away understanding. So you have time to think and strategize, and that only works if your SDLC allows that.

Cheers,

Diego Pacheco

Popular posts from this blog

Cool Retro Terminal

Harness Engineering

AI coding Agents Evolution