The AI Scarcity Mindset
What does it mean to view AI compute as a limited, rather than abundant, resource?
I started my journey in generative AI when ChatGPT was first launched. It was a time of great promise, but also scarcity.
AI models (I was using GPT 3.5 and 4.0 at the time), were very context limited and had narrow reasoning capabilities.
I was working on early versions of retrieval augmented generation (it wasn’t called that at the time), where I’d provide the model with content from a data store. I also manually set up tools to deliver content from the Web to the model via Python scripts.
This work was extremely frustrating. Not only would the model hallucinate a lot, but you could only put so much information into its context. I had to put content through customized summarization pipelines and carefully construct my prompts to make sure the LLM saw the right things.
This was the minimum required to allow these models to function. There was also a financial incentive because input and output tokens were, on a per-token basis, expensive to generate.
The Scarcity Mindset
These early experiences shape how I use generative AI today. Even though models have larger context windows and can reason much more effectively, I’m still stingy with it. I don’t view inference as an unlimited resource and am careful with how I use these models.
This scarcity mindset influenced how I looked at a recent Anthropic study, focusing on AI’s use for work tasks. Anthropic found that AI is helping people engage in tasks that deliver more economic value over time. This is fantastic, but I thought: how much is this actually costing us?
So I decided to measure it. The result of this work is the Secrets of the LLM Whisperer study, which looked at the habits and behaviors that contribute to inefficient (and more expensive) LLM use.
A key driving factor is that, although we have more AI inference available to us, it is still a scarce resource that is being heavily subsidized. AI labs are feverishly buying data centers to meet the voracious appetite for inference. And, a backlash is growing, with some states imposing moratoriums on data center construction.
The cost of inference is also moving up. Anthropic is putting its most powerful current model Fable, behind stringent usage caps and pay-as-you-go payment schemes. Many are angry about this, but labs are heavily incentivized to stop subsidizing AI inference so heavily and have users pay closer to the market rate for it.
The study results and these market forces inspired me to launch this newsletter. In it, I’ll explore questions like:
How can we use AI more smartly?
What are the strategies and tactics being used to conserve resources (such as water, land and energy) consumed by AI?
How can we make the social and economic costs of AI lower so ROI is higher?
How can we operate AI more securely (which has its own efficiency benefits)?
Thanks for joining me on this journey.



