Intelligence Too Cheap to Meter
What happens to GTM as the cost of tokens goes down?
Here in Austin, we’re in the sweaty denouement of summer. Family vacations are over, the humidity is up and every school-age kid is trying in vain to slow the passage of time as the start of school looms.
Despite the heat, nothing about the summer of 2026 has been slow at the intersection of GTM and AI. Deals are getting done (even if they’re “a kick in the gut”). Meanwhile, every AI model provider is trying to outdo the other on the “committing a cybersecurity felony” leaderboard.
It’s been eventful for me personally as well. Earlier this week, I announced a new forward-deployed engineering1 agency called The Agent Deployment Company (ADC) alongside a fantastic team of engineers. Nothing’s changing with my role at Gradient Works, but I’m excited to play a part in helping GTM teams build the technical foundations they need to be successful with agentic AI.
Like everyone, I’m just trying to keep up with the constant change, wrestle with where all this is going, and make my bets—like ADC—along the way. One area I’ve been thinking a lot about is token costs and what they mean for GTM.
Tokens aren’t really a measure of intelligence, but they’re a proxy for how today’s LLMs process information. If tokens get cheaper, it’s approximately the same as saying that intelligence is getting cheaper.
In the rest of this article I’ll share where token prices seem to be heading (spoiler: it’s down). Then I’ll look at a couple of follow-on effects we’re already seeing in GTM: how the Slop Wars are making scalable demand gen so hard and how one ops team is reorganizing itself to build for internal customers in a way that would have been ludicrous a few months ago.
Will tokens become too cheap to meter?
In 1954, Lewis Strauss, Chairman of the United States Atomic Energy Commission (and Oppenheimer villain), told us that energy would one day be “too cheap to meter”. That obviously didn’t come to pass, but it might have if a few things had worked out differently.
AI is likely to be as transformative as harnessing the atom, and we’re yet again trying to predict the economic impact. One thing people can’t decide on is whether it’ll be incredibly expensive or incredibly cheap. We’ve recently had a few points in the “incredibly cheap” column—at least on a per-token basis.
OpenAI last week dramatically cut the cost of the Luna and Tera flavors of their GPT-5.6 model. That makes them attractive in terms of price vs performance, as you can sorta see in this hard-to-read chart:

This is part of a longer trend. On Monday, The New York Times published an article on “tokenomics” that linked to a working paper from economists at the National Bureau of Economic Research. The authors analyzed 380 trillion tokens of AI consumption across 400 models from an OpenRouter dataset.
They found that, while agents are causing token consumption to skyrocket, the price per token has declined precipitously over the last 18 months.

There are several reasons for this decline, but the authors note that platforms charge less for “cached” tokens that are sent to the model multiple times during long-running agent sessions.2 They also suspect that open source models may be having an impact as well. And all of this was already happening before OpenAI started cutting prices on frontier models.
Benedict Evans recently published a fantastic piece on this topic. He used to be a telecom analyst and that turns out to be relevant to the AI world. One of his most striking observations is comparing tokens to bytes:
That makes mobile data a more fruitful comparison here. Mobile networks have marginal cost for capacity, and like AI they had an enormous surge in usage 15 years ago, that overwhelmed capacity and had carriers scrambling to add capacity and rebalance their pricing. Meanwhile, selling bits looks superficially similar to selling tokens: it’s an opaque measure of marginal cost that doesn’t map in any transparent or intuitive way to use cases or value, and needs to be replaced with bundles of some kind.
Those of us who lived through the early era of mobile data plans will recall how frustrating it was. How many bytes would it cost me to visit this web page? Who knows! Tokens are the same, but probably even more disconnected from any intuitive reality.
The per-token cost of AI looks likely to continue declining. There’s so much competition across platforms and switching costs are so low that it’s hard to see how anyone can get enough differentiation to drive up prices.
We’re already on our way to bundling, as Evans suggests. OpenAI and Anthropic’s subscription plans are already bundles—even though they’re essentially just packaging cheaper blocks of tokens in exchange for more predictable subscription revenue. That’s a strong step in the direction of “too cheap to meter”.
Of course, cheaper tokens don’t mean AI spend overall will decline. We’ll likely consume more tokens, thanks to the Jevons Paradox of it all. That said, anything that transitions away from an opaque, unpredictable consumption metric towards more predictable value-driven pricing will help us better manage COGS and gross margin.
Slop Wars: Episode VI - Return of the Mods
Ever-cheaper AI combined with the incentive to capture attention means there’s no escape from AI—anywhere.
I don’t know about you, but I’ve found LinkedIn increasingly unbearable the last few months. My feed has become a stream of algorithmically selected posts from strangers, each bearing the distinct sheen of whatever it is Claude sounds like these days.
Well, perhaps the tide is turning. Last Friday, LinkedIn finally fired their first shot in the Slop Wars, adding an actual honest-to-god “Seems like AI slop” menu item. I’ve been pressing it liberally. Hopefully we can reclaim LinkedIn for human-driven cringe combined with some—occasionally useful—GTM advice.
Obviously LinkedIn is just one front in the Slop Wars. Another front involves the increasingly important AI cousin of SEO, GEO (or AEO). LLMs really like Reddit, which means that subreddit mods now find themselves fending off armies of bots looking to inject certain brands into the conversation.
The Verge has a great article about this phenomenon that includes the story of a moderator in r/SkincareAddiction. It’s an increasingly tricky job:
Years ago, moderators would look for certain telltale signs of spam or covert promotional content: a review that was a bit too glowing, product links with tracking tags, or before and after pictures that were the same as those included in product listings for the brand. AI search has changed how marketers attempt to game systems such that some SEO experts say that brands don’t even need a backlink anymore to effectively get their websites cited by a chatbot — merely a mention is enough. A suspected spammer might pose a question like “How does everyone keep track of their skincare routine?” that at first glance appears to be authentic engagement, but then a wave of comments will mention AI skincare apps, for example.
The obvious outcome of this is that LLMs are increasingly training on, and searching for, content created by other LLMs. Dead internet theory, indeed.
Any “scalable” GTM acquisition channel that’s not purely pay-to-play will ultimately succumb to the tragedy of the commons. That said, you’ve got to try to find your GTM alpha where you can, so you might as well invest in GEO while it could give you a slight edge. Just try to be reasonably ethical about it.
Just remember, the only demand gen channel you can invest in long term is good old-fashioned unscalable human effort.
Building an ops team of builders
Earlier this week, over happy hour drinks with a friend, I got a demo of an awesome SaaS product for CS teams. It integrated account scores, task management, calendaring and customer communication into a single streamlined workflow. There was AI chat for getting quick answers across all that data. Everything synced back to Salesforce. It even had a nice dark mode and achievement badges.
Except it wasn’t a SaaS product. It was a bespoke internal app built by their ops team. And they built it in under two weeks. Now it’s in production.
I’m not credulous enough to believe this is magic. There will be maintenance and bug fixes. There are real costs and tradeoffs associated with going down this route. But six months ago, nobody would have even considered attempting to build something like this. It just wasn’t a thing. Now? It’s a reasonably credible option. Why? What previously would have been person-years of human effort simply boils down to—you guessed it—token costs.
My friend doing the demo is VP of Operations at an 800-person company. He’s made a big strategic bet on this kind of internal build capability. I found the way he’s organized his team around this fascinating.
He’s divided his org into squads—each of which is essentially a small product team capable of doing everything from PM to engineering. The squads are assigned to customers within the business (e.g. new business, CS).
Squads have a matrix reporting structure. Some members of a squad are on the “core ops” team and report directly to the squad leader. Other members of the squad have functional roles (e.g. RevTech, Analytics) and report to their functional leader but have a dotted line to the squad leader. Most of their time is devoted to the squad. This attempts to balance squad work with “platform” work to make sure there’s knowledge transfer across squads.
One of the big challenges for ops teams is getting so backlogged with “keep the lights on” (KTLO) work that they can’t actually build anything new. One way he’s addressed this is to challenge the teams to take two weeks out of every quarter to set the KTLO work aside and focus solely on building. The product I mentioned above is the output of one of those two-week sprints.
There’s a lot of balancing going on here, but he’s been at this for a few months now and it’s working well so far. It’s a bold strategy and I’m eager to see if it pays off. This team design—augmented by agents, of course—may just be the right way to build at the surface.
Wrapping up
As token prices go down, the cost of gaming every possible demand gen channel with slop goes down—ruining the commons for all of us. At the same time, it empowers teams that fully embrace this new reality to build in ways that would have seemed literally crazy six months ago.
These are just two relatively small areas of GTM, but they’re representative of the larger change that’s coming. The marginal cost of human-like (if not always human-level) intelligence is dropping precipitously. We’re only starting to feel the effects.
I’ve talked trash about FDEs in the past. I’ve since come around. Meaningful agent deployment is blocked by hard technical problems related to access, skills and context. These problems are deep enough and different enough across organizations that they need real engineering skill to tackle. I don’t really want to hand it to Palantir, but they’re right about this model—for this particular moment in time.
OpenAI, for example, charges 1/10th the price for cached tokens vs non-cached.



