Cognatu
Insights

The Token Is the New Bandwidth

Why AI's economics look like 1999, and what becomes valuable when intelligence gets cheap

By Shirish Garg  ·  August 15, 2026  ·  6 min read

Over the past few months I went through 33 AI products, one at a time, trying to work out how each of them planned to make money. Most had no answer. A couple had subscriptions converting in the low single digits and were calling it a plan.

The four companies supplying the infrastructure those products run on will spend about three quarters of a trillion dollars this year5.

That mismatch is not new. Every technology wave turns something expensive into a commodity, then has to work out who pays for it. The internet's commodity was bandwidth. AI's is the token: the unit every model's output gets billed in.

Token economics are running the internet's script almost beat for beat: the price collapse, the overbuild, the business model nobody has invented yet. The ending is the part worth retelling, because it wasn't technical. It was advertising, reinvented, and it built the two most valuable media companies in history.

Act one: the price collapse

OpenAI opened the GPT-3 API in June 2020, behind a waitlist it did not drop until November 20212. A million Davinci tokens cost $60 then, and stayed there until a price cut the following September3. Token prices across the inference market have fallen roughly 600-fold between 2020 and 2026, with economy-tier models on a price half-life of about thirteen months, comfortably faster than Moore's Law1. One research group puts the halving time of inference cost at roughly every 2.6 months5. Either way the decline is steeper than compute during the PC revolution, or bandwidth during the dot-com boom4.

The price of a million tokens on a log scale, falling from $60 to $0.10 between 2020 and 2026, roughly a 600-fold decline. A ghosted line shows bandwidth costs falling in the same shape between 1995 and 2003.
Economy-tier inference, on a log scale. The dotted line is bandwidth, 1995–2003.

This is precisely what happened to the pipe in the late 1990s: telecoms overbuilt fibre against forecasts that never arrived, and bandwidth became nearly free. Wonderful for the internet. Ruinous for anyone whose business was selling the pipe.

Cheap tokens are doing the same for intelligence. A product that couldn't afford to spend a dollar answering a question can now afford to answer a thousand. AI stops being a premium feature and becomes infrastructure.

But a collapsing unit price carries a warning the industry doesn't like to say out loud: when the commodity gets cheap, the commodity stops being the business.

Act two: spending like it's 1999

The buildout, meanwhile, has gone vertical. Amazon, Google, Microsoft and Meta guided to roughly $725 billion in combined 2026 capital expenditure on their Q1 2026 calls, up about 77% from around $410 billion in 2025, and two of them have raised since6. That is a step change in how much of the corporate cash flow of the largest companies on earth is pointed at one bet.

Against it, the revenue is thin. Sequoia's David Cahn put the annual gap between AI infrastructure spend and AI ecosystem revenue at roughly $600 billion back in mid-2024, and capex has accelerated faster than monetisation ever since7. Allianz Research now measures the divergence between AI capital spending and revenue growth at around 46%, already beyond the 32% seen in the 2001 telecom cycle that preceded a brutal multi-year correction8.

Those 33 products are the other end of that number. Three quarters of a trillion dollars of supply, and a demand side that mostly cannot answer the question of how it earns a living.

None of this means the buildout is wasted. Demand for intelligence, like demand for bandwidth, will eventually absorb everything being built and ask for more. What's missing now is exactly what was missing in 2000: a demand-side business model that lets ordinary economic activity flow through the new pipes at scale.

The internet found one. It took two attempts, and almost nobody remembers the first.

Act three: what actually paid for the internet

It wasn't saved by subscriptions, and it wasn't saved by e-commerce margins. It was saved by advertising.

The first attempt failed. In 1994 AT&T paid $30,000 for three months of banner space on HotWired and got a 44% click-through rate9. That number started a bubble, venture capital poured in behind it, and the dot-com bust arrested the whole thing10. Selling attention by the pixel could not carry an industry.

What worked was intent. If someone types "best running shoes," you already know what they want. An ad for running shoes isn't an interruption; it's almost an answer. Google launched AdWords on that insight in October 2000, promising ads that were relevant, clearly separated from the results, and free of the pop-ups and animation everyone else was selling11. It took another sixteen months to find the pricing that made it a machine: AdWords Select, in February 2002, charged advertisers only when someone actually clicked, and ranked ads by click-through rate as well as bid, so a more useful ad could outrank a richer one12. The innovation was never "put ads on the internet." It was: put the commercial message next to the user's intent.

By mid-2004, the quarter Google went public, US internet ad revenue hit $2.37 billion in a single quarter. That beat the dot-com era's peak13. A decade later Facebook ran the same play on a different scarce signal, social context: what you like, who you follow. It still earns 97.6% of its revenue from ads14.

So the pattern is: a new medium makes something scarce, and whoever monetises that natively takes the wave. Search ads for search. Feed ads for feeds. Never a banner bolted onto the new thing.

Three waves compared: search made pages abundant and intent scarce, monetised by sponsored results priced per click, and Google won. Social feeds made posts abundant and social context scarce, monetised by in-feed ads, and Meta won. AI assistants are making tokens abundant and intent scarce again; the native format is not invented yet.
The winner is never the one who bolts the old format onto the new medium.

Act four: what's scarce in AI?

Now run it forward. Tokens are becoming abundant. Compute is getting cheaper. Models are increasingly interchangeable. What's left that's scarce?

Intent. And AI conversations contain more of it, stated more explicitly, than search ever did.

Someone asking an AI "I'm moving to Bangalore. Which neighbourhood should I live in?" has expressed something enormously valuable. So has "I need a laptop for video editing under ₹1 lakh," and "How do I get health insurance for my parents?" Nobody typed those into a search box as keywords; they said them, in full sentences, with context attached.

But here the parallel to search breaks, and the danger starts. Search engines mostly know what you're looking for. AI assistants increasingly know why: that you're worried about money, planning a divorce, job-hunting behind your employer's back, trying to conceive. Building the ad business by quietly compiling all of that into permanent server-side profiles wouldn't be a business model. It would be the internet's original sin, re-committed with a far more intimate data stream. A scandal on a delay.

Intent without identity

There is another way to run the play, and it's the one the technology of 2026 makes possible in a way 2000's didn't.

The advertiser doesn't need to know you. It needs to know that someone, right now, is in the market for a product like theirs. That signal is just the commercial category of the conversation. It can be extracted on the device and matched on the device, and it never has to leave. The conversation stays put. So does the profile, and so does the identity. The intent gets monetised. The person stays unknown.

A conversation about moving to Bangalore stays inside the device, where a small model reduces it to a commercial category. Only a two-field payload (category real-estate, intent high) crosses the boundary to the auction, which returns one labelled ad. The conversation, the profile, and the identity never leave.
Advertisers bid on the category, because there is no person to bid on.

That's the native format the pattern demands: advertising built next to intent, the way AdWords was, but without the surveillance apparatus the last wave bolted underneath it. This time the matching can happen where the data already lives.

That is the bet Cognatu is built on, so treat me as an interested party. I could be wrong about the timing, and I do not know what the format finally looks like; the last two waves both took a second attempt to find it. I am fairly confident about the direction.

The internet made the pipe cheap, and the companies that monetised intent became Google and Meta. AI is making intelligence cheap. It will create companies of the same scale. The open question is the one the last generation answered carelessly: whether they get built by making the user the product, or by proving that they don't have to be.

If you're building one of these

Don't take the last section on faith

The demo runs a real conversation and matches an ad on the device, with nothing but a topic category leaving it. Open your network tab while it happens. The whole argument is that there's nothing in there worth reading. No account needed.

Watch it run Or read how the SDK works →

Next: Your Life Shouldn't Have to Log In. The same primitive, pointed at personal context instead of advertising.


Sources

  1. Mingdeng Du, "Tiered Super-Moore's Law: Price Evolution, Production Frontiers, and Market Competition in Large Language Model Inference Services", arXiv:2603.28576, 30 March 2026. Documents the ~600-fold decline across 2020–2026; economy-tier price half-life 1.10 years, mid-tier 1.55 years.
  2. OpenAI announced the API on 11 June 2020, in a waitlisted private beta, and removed the waitlist on 18 November 2021.
  3. OpenAI's pricing page, archived 4 December 2021: Davinci $0.0600 per 1K tokens, i.e. $60 per million. Cut to $0.0200 on 1 September 2022.
  4. Guido Appenzeller, "Welcome to LLMflation", Andreessen Horowitz, 12 November 2024. Source of the comparison to PC-era compute and dot-com bandwidth.
  5. Xiao et al., "Densing Law of LLMs", arXiv:2412.04315, §3.4: "the inference costs for LLMs halve approximately every 2.6 months."
  6. Q1 2026 guidance from the four companies' own investor relations, e.g. Meta's Q4/FY2025 results, as tallied by the Financial Times and reported in Tom's Hardware. Note: Meta has since raised its 2026 range to $125–145bn, so the combined figure is now higher.
  7. David Cahn, "AI's $600B Question", Sequoia Capital, 20 June 2024.
  8. Allianz Research, "AI capex cycle: war-proof for now" (Subran, Utermöhl, Dejean and Hirt), 25 March 2026. Source of the ~46% capex-to-sales growth gap against 32% in the 2001 telecom cycle.
  9. Digiday, "How the Banner Ad Was Born". The $30,000 three-month placement is attributed to HotWired CEO Andrew Anker; the 44% click-through rate is GM O'Connell's recollection, and does not appear in the participants' own oral history.
  10. Digiday, "A History of Ad Tech, Chapter 2: The Ad Net's Golden Age", 2023.
  11. Google, "Google Launches Self-Service Advertising Program", 23 October 2000. Note the launch pricing was CPM: "$15 or 1.5 cents an impression" for the top position.
  12. Google, "Google Introduces New Pricing for Popular Self-Service Online Advertising Program", 20 February 2002. AdWords Select, cost-per-click pricing, ranked by click-through rate as well as bid.
  13. IAB / PricewaterhouseCoopers, Internet Advertising Revenue Report, full year 2004, which records the dot-com-era quarterly peak of $2,123m in Q4 2000; Q2 2004's $2.37bn is given in the Q2 2005 report.
  14. Meta Platforms, Form 10-K, FY2025: advertising revenue $196,175m of $200,966m total.

Cognatu builds privacy-first infrastructure for AI apps · cognatu.com