A bottom-up model of AI demand

One way to model AI demand is to go top down and start with compute supply. If you assume users will absorb every watt of compute that providers can make available, the supply forecast becomes the demand forecast.

But I find that unsatisfying as a forecast of demand, even though I don’t think we’re anywhere near a compute glut. I want to understand how people are actually using tokens today, and what will make them use even more tokens.

I prefer a more bottom up approach to modeling AI demand. One that is driven by usage and behavior. I've been trying to answer a series of questions along these lines. What do usage patterns for various knowledge workers look like today? What is driving them? How are they likely to trend?

The easiest place to start is with myself, so let me walk you through what my AI usage looks like.

To get the data, I exported my token usage analytics from both the OpenAI and Anthropic consoles and added my local Codex and Claude Code usage. I've charted the result below which understates things a bit since it doesn't include my LM Studio, Gemini or OpenRouter usage. But, it is directionally correct.

Over 12 months, my usage goes from about 8 million tokens a month to about 300 million, roughly 40×. This probably looks similar to your usage chart if you consider yourself an early adopter. What I find most interesting, though, is that my usage growth closely mirrors key model capability jumps and harness improvements.

I'd break down the various inflection points in my usage as follows:

  • Prehistoric times: I was still mostly doing Q&A and informational tasks, plus the occasional deep-research-style prompt.
  • December 2025. My usage spiked after the release of Opus 4.5 in November, when agentic tool calling got really good; my Anthropic usage roughly quintupled in a month. This is also when I started using Claude Code heavily for non-engineering tasks: organizing folders, managing my growing library of film photos, etc.
  • February–March 2026. I started using the Codex app heavily after the Feb launch. The sub-agent paradigm is suddenly everywhere, and boy does it use up a lot of tokens. GPT-5.4 in March pushed things further; my OpenAI usage more than doubled in both February and March.
  • June 2026. After the late-May launch of Opus 4.8, my Anthropic usage more than doubled. I got really comfortable with parallel agents and after my egobench project, started investing in personal benchmarks, and per-project evals and guardrails. I think this made me quite comfortable reviewing work at a higher level of abstraction and default to delegating larger chunks of work to both Opus and GPT-5.4. It's also worth noting that there is some built in reflexive demand as quite a few of my project guardrails are just skills that themselves use up a lot of tokens in a LLM-as-judge paradigm.
  • July 2026. This is when scheduled tasks really clicked for me. Codex automations shipped with the app in February, and Claude Code got desktop scheduled tasks in early March, but I didn't start using them until mid-June. Once I was comfortable reviewing at a higher level of abstraction, handing off unattended runs was the natural next step. July was the first full month of it, and my Anthropic usage more than doubled again. I now have a ton of scheduled tasks I review daily and weekly.

I think of my token usage as being driven by two axes: (1) model capability increases, and (2) harness improvements. Both of these combine to drive new and sticky behaviors that use more and more tokens.

I'm an early adopter, not a typical user. So why should an n=1 usage chart tell you anything about broader AI demand?

I think the answer is that AI diffusion has a lag and we’re incredibly early in adoption. The behaviors I've picked up over the past year (delegation, parallel agents, unattended scheduled runs) will reach other knowledge workers soon enough, just with a lag.

If that's right, the interesting questions for forecasting AI demand aren't really about compute supply. They're about how quickly behaviors like these spread, and how many more tokens each one burns once it does. Judging by my last year, both are still going up. It's hard to predict what other new behaviors will stick for me, but I'm sure they will drive even more token usage.