I have been on Substack for years, but only recently started to devote more time to the platform. Like many of you, I’m keenly interested in learning how to reach my audience and grow on Substack. So, I began by analyzing dozens of Substack accounts for clues.
Very quickly, I noticed something strange. Some newsletters were growing very quickly, gaining 1,000s of subscribers in 30 days. Other people have been on the platform publishing regularly for months, but have seen very slow or limited newsletter growth.
I wondered why so I started reading advice from others (mainly marketers) about how to grow on Substack. They often pointed to their own subscriber acquisition success as evidence their recommended growth strategies were effective.
I found their takes interesting and useful, but what they all had in common was that they were all in the ‘how to market on Substack’ niche. There’s a natural audience for this type of content on Substack. Everyone here wants attention, so it’s not surprising many in that area are able to grow their accounts quickly.
I’m not in the Substack marketing niche, so I looked for other examples of high-growth accounts. My survey wasn’t scientific, but I noticed that people in other niches on Substack had been on the platform for quite some time, often years. When these people post notes, they get lots of engagement (because their audience is larger) and seem to benefit from a growth flywheel.
However, these people aren’t the norm. Most have much lower subscriber counts. This is not necessarily a bad thing. Many people are not on Substack to grow exponentially, but to find a place where they can simply share their perspectives and find a community.
But, if we’re being honest, many people on Substack would like to earn from their writing. Earning requires growing an audience. That’s the promise of Substack after all, right? So I decided to dig deeper into the Substack recommendation engine to learn more about how it operates, surfaces content and promotes account growth.
As an aside, I really wanted to find out why those ‘Do Your Thing Substack’ posts seem to work so well at attracting subscribers. I mean these posts where people ask others to connect around their topic can generate hundreds of comments and thousands of likes almost immediately. That type of post might not be your cup of tea, but you may be curious about why they appear to be effective.
My goal isn’t to give you a “proven blueprint” to go from 0 to 1,000 subscribers in 30 days. Instead, the purpose of the Doing AI Efficiently Operating System (this newsletter is at the system’s center) is to provide you with first principles insights and education on AI so that you can use it more effectively.
At its core the Substack recommendation engine is AI (my acronym for it is SAIE). I want to help you understand SAIE so you can improve your approach to the platform and attract more people to your insights and writing.
There are no guaranteed shortcuts to gaining an audience on Substack. However, armed with the right knowledge, you’ll be able to better calibrate your activities on the platform to set yourself up for success more effectively.
Contents: A Guide to the Substack AI Recommendation Engine
Four Questions I’ll Answer
Fortunately, Substack has provided all the information we need to understand how its AI recommendation engine works. There are other guides to the Substack Notes algorithm. However, the popular ones I’ve read aren’t grounded in the technical details of the system. There’s a lot to be learned once you read research papers highlighted by the Substack team and the team’s statements about SAIE.
I’ll provide you with an overview of the system’s mechanics, as I understand them, at a level appropriate for non-technical readers. Note that I haven’t spoken to Substack about this essay, and I don’t claim to have insider knowledge of how the system works. My analysis is based solely on review of publicly available information.
I’ll answer four questions in this essay:
What is the Substack recommendation engine designed to do?
How does SAIE work?
Why is SAIE pumping harder for some people versus others (and does it matter)?
Why the house always wins: Substack is optimized for subscriber acquisition, not retention (that’s on you)
If you’re here for the first time, thanks for visiting. Please subscribe for additional first principles analysis, strategy guides, courses and education on AI systems, in generative AI and beyond.
1. What is the Substack Recommendation Engine Designed to Do?
Let’s not forget that Substack is a business. Its goal is revenue maximization. That means, like other platforms, it needs people to:
Discover the platform
Find and subscribe to publications
Pay for newsletters (this is how it makes money)
One of the reasons I didn’t try to grow my newsletter on Substack in the past was the discoverability issue. If I set up a newsletter on Substack, I still had to find readers by promoting my content on other channels like LinkedIn. If that were the case, I thought, why not just use LinkedIn, or send people to my own platform? Substack was a hard no for me.
When I came back years later, I had the same attitude. Then I discovered the Substack feed. This changed the game for me because I thought: “Okay, Substack is trying to solve the discoverability issue by providing me with a means of reaching others on the platform.” That seemed worth the effort to me, so I decided to invest time here.
So, I don’t mind that Substack is optimizing for discoverability and paid subscriptions. It’s a win-win, right? Almost.
There’s a catch. The SAIE is not optimized for your success as an individual publisher. Substack wins when publishers, in the aggregate, do well on the platform. Every subscription counts equally. The system doesn’t care whether people subscribe to YOUR publication, only that people subscribe to A publication.
The other things SAIE is optimized for are user data processing, rapid analysis and predictive power. For reasons that I’ll explain later, the system needs a steady stream of platform activity data generated by you. Only show up a couple of times a month to post an article and send out a Note announcing it? Guaranteed audience growth failure mode.
SAIE optimizes for signal accumulation. Every click, like, restack, subscription and payment feeds the system and influences what you’ll see next (and whether your account will grow). I noticed the engine pushing content to me from the beginning. I clicked on some of those “get 1,000 subscribers in 30 days” articles, and that’s all I saw in my feed for a couple of weeks. This changed, but only when I adjusted my reading, following and engagement habits.
SAIE can help you gain subscribers, but retention is on you. Substack only looses if users leave the platform and don’t pay for subscriptions. SAIE is agnostic about whether it’s your content, or someone else’s, that’s recommended.
Understand what SAIE optimizes for and you’ll recognize how to use the system to your maximum advantage.
2. How Does SAIE Work?
SAIE is a complex system. While I didn’t receive a briefing from Substack about its inner workings, I found several valuable sources of information that helped me to learn about SAIE’s technical foundations:
Mike Cohen, Substack’s Head of AI, wrote this article in October 2025 outlining the system’s technical architecture
Amanda Bray, of the Publishing Spectrum conducted an interview with Cohen in late July. In the interview Cohen confirmed that it is still using the methods outlined in his October article to run the recommendation engine
This article, also published in the Substack official newsletter in October 2025, provides specific information about the user data signals Substack is using in the engine
From User Activities to Recommendations
At its core SAIE uses machine learning, specifically sequential models. A sequential model is designed to handle ordered or time series data, where the “order or sequence of the input matters”. Sequential models are used heavily in recommendation engines. For example, in order to predict what content a person might click on next, information is gathered about their previous activities over time. Consider this question: If a female user, age 30, who has used a platform for 6 months, views A, B and C content over a specific time period, what type of information will they be most likely to click on next? A sequential model might answer: “This user will most likely click on an article showing a picture of a dog.”
For Substack a properly trained and calibrated sequential model is valuable for helping to predict what you’ll:
Read
Interact with (like, comment, restack, share)
Subscribe to
Pay for in the future (the most important success metric)
Have you ever noticed how much data Substack collects about your platform use? In the Substack dashboard can see exactly how many times your Notes were viewed, whether a view resulted in a subscription, how often your readers click on links in your emails and many more pieces of information. This is the type of data that’s fed into SAIE to help it predict the content users will be most interested in viewing, and what should be recommended to them.
Sequential models owe some of their architecture to large language models (LLMs). As I explained in my essay on AI text watermarking, an LLM translates strings of text into numbers so that the model can process it.
Sequential models also translate information. In order for the model to understand user data it must be converted into embeddings (lists of numbers) representing different on-platform events. Embeddings are very important because they allow the model to rapidly reason over numerous parameters and make associations between data types more easily. As Cohen said in his essay describing the system:
“[W]e’ll integrate ... sequential embeddings into ranking ... The sequential approach brings two major advantages: first, the final ordering will be tuned to the flow and momentum of your current session, understanding which posts make sense as the next step in your reading journey. But perhaps more importantly, the user representation itself will be much smarter, incorporating attention mechanisms and 10x more features than before, making the ranker better at determining what posts are good matches for you in general, not just in this moment.”
The types of data used by the predictive model can vary, but may include:
Where you’re located
What language you speak
Which publications you’re subscribed to
Which creators you follow
What interests you’ve specified during Substack onboarding
How often you’ve liked content
What you’ve restacked
Sentiment analysis of content you post
Activity Time Periods, Signal Strength and Negative Actions
There are three factors that may have a big impact on SAIE’s recommendations.
Time period: The system considers what you’re interested in reading now and in the past, and, as Cohen notes tracks “the momentum and direction of your current session.” Importantly, the system doesn’t just favor your most recent interactions, but develops a sense of how you use the platform over the long-term. Cohen said:
“[W]e still preserve your long-term reader embedding alongside the sequential state ... So even if you’re deep in a poetry session, the system still knows about your enduring love of sports journalism, your subscriptions to tech newsletters, your history of engaging with climate content. Those long-term signals ensure that other parts of your interest graph stay in circulation.”
Signal Strength: The more information the system has about you, the better it can optimize for your interests and reading habits. At the same time, a rich data set about your writing enables the system to better recommend YOUR content to others.
Negative Actions: Cohen didn’t discuss this in his article (or his other public statements), but it’s also critical to understand that sequential model-powered systems can be designed to track what I call negative actions. These are signals you send to the system when DON’T do certain things.
For example, let’s say the system is recommending a certain type of content, but you don’t interact with it. SAIE learns this content isn’t worth showing to you. This may be why certain Notes get buried, while others, like the infamous ‘Do Your Thing Substack,’ posts are more likley to appear in your feed. You may not like them, but the system has been taught that they get engagement AND drive subscriptions.
Pinterest, which helped to advance the state of the art in sequential models had this to say about using negative actions to fine tune recommendations: “The second, impression-based negative sampling, selects Pins ... that the current user has viewed but not engaged with further, suggesting low interest. Our findings ... reveal that impression-based negative samples are more effective ... [and improves] ranking performance.”
Cataloging Key Substack Signals
I spent a lot of time in the last section discussing signals Substack uses to identify what types of content to show you in its feed (and recommend to others). That’s because it’s the single most important factor that determines your fate on the platform.
Provide the right data and signals to Substack and, given enough time, it can help your newsletter grow.
Below I’ve provided a catalog of the types of signals Subatack is may be feeding into SAIE to help you focus and optimize your activities on the platform.
Remember: It’s about signal strength and time. Think of it in these terms (this isn’t how Substack does it, but it helps explain the concept):
Signal Density is the average volume of topical and engagement activity data you have generated.
Time is how long the system has had to collect data about you, your content and actions.
The more you use Substack, and the more time you spend on the platform, the better the odds the system can:
Recommend content to you
Put YOUR content in front of people who might subscribe to your newsletter
Although there may be a negative action penalty, it might be used to help Substack better understand who your ideal reader is over time.
Your Activity Isn’t Wasted Effort
Those Notes you send out that get no reply? They’re not a waste of time. They are valuable because SAIE is able to use this data to better inform its predictions and recommendations. And, your content is used to enrich the feed.
Have you ever seen those Notes published weeks ago in your feed? That’s SAIE at work, using its available inventory of content, whether it was published yesterday, or five months ago, to help keep you engaged.
Below is a listing of signals SAIE may be using to inform content recommendations.
The system is designed to accumulate as much of this data as possible. The more signals you generate, the better SAIE can understand who you are, what you write about, and who might subscribe to your work.
3. Why Is SAIE Pumping Harder for Some People Versus Others (And Does It Matter)?
Let’s address to the elephant in the room: Why are some people gaining subscribers faster than others? There’s no secret formula to success, but understanding SAIE provides some useful insights about what the system favors. And, like all algorithmic systems, it has its weaknesses.
Understanding SAIE Content Bias
Does SAIE censor content, individuals or accounts? Yes. Every recommendation system like the one Substack runs has safety and spam filters.
Given Substack’s investment in Pangram for potential AI content detection, it’s unclear whether content flags are used to de-emphasize certain types of posts. Pangram scanning isn’t activated automatically, so even if is a signal, it’s likely weak and secondary. This could change if automatic scanning is enabled.
Importantly, there may be an inherent biases in the system that can fuel the perception the algorithm is unfair.
First, recommendation systems are known to potentially suffer from “popularity bias.” Specifically in this paper by Anastasiia Klimashevskaia and colleagues the authors suggest: “algorithms may have a tendency to focus on already popular items in their recommendations. As a result, the already popular (“Blockbuster”) items ... receive even more exposure through the recommendations, which can ultimately lead to a feedback loop where the ‘rich get richer’”.
In addition to popularity bias, the recommendation engine may work better for users with a richer data history. In an analysis Shengyu Zhang and colleagues note: “Recommendation performance usually exhibits a long-tail distribution over users — a small portion of head users enjoy much more accurate recommendation services than the others. [There are] two sources of this performance heterogeneity problem: the uneven distribution of historical interactions (a natural source); and the biased training of recommender models (a model source).”
What this research suggests is that:
Users with a longer account history may have their content recommended by the system more often versus those with a less user data (uneven distribution)
The model may be trained on data that causes it to favor more popular topics
Other factors that may further bias SAIE include:
Revenue Optimization: SAIE optimizes for subscription events, not content quality. Content that reliably generates subscriptions gets favored.
Topic Signal Density: Topics with a lot of data associated with them (clicks, restacks, subscriptions) can deliver higher quality predictions. Are you writing about marketing on Substack? Your personal journey that netted you six figures in subscription volume? These topics are popular and you may benefit from a positive feedback loop of attention and subscriptions.
But, if your topic is more niche, like food safety, the history of hats, or outside of a popular topic area, the system has less information to go on, and may surface your content less (and your growth may be slower)
A biased recommendation engine:
Cultivates a monoculture that rewards only certain types of writers and topics
Results in a lack of support and visibility for newer accounts. If the rich get richer affect is pronounced, the Substack user base may become disillusioned, and the community will be less attractive to new users; in this situation, growth stalls
“Hacking” the Algorithm
There are some early signs that some are either purposefully or accidentally taking advantage of some of the weaknesses of recommendation engines. Some have figured out ways to ‘hack’ SAIE to generate high account growth in a short period of time. Two of these strategies are explained in the images below:
“Do Your Thing Substack” Notes: Discussed previously, this strategy takes advantage of SAIE’s preference for positive content engagement signals to generate outsized visibility for certain accounts. While not all Do Your Thing Substack posts are successful, they reward stye over substance. Fast growing accounts with thin content are outperforming new accounts where authors are developing in-depth analysis, opinion and other rich content.
“Subscribe for Subscribe”: Users are engaging in mutal subscription exchanges. This strategy directly targets the engine’s primary recommendation driver, subscriptions, to drive engagement and accelerate account growth velocity.
Substack was founded to reward writers seeking a home to share substantive content with audiences. While the SAIE has boosted Substack’s userbase, the engine could also become a major source of user frustration and distrust if it is not successfully spreading the wealth across the userbase, and rewarding high-value, high-signal content creation.
4. Why the House Always Wins
Substack wins when you win. But, SAIE’s goal is to drive revenue for the company, not a specific newsletter.
Here’s the lifecycle the engine may optimize for:
New user joins on Substack through a newsletter, referral or other means
SAIE shows them relevant content from the feed
They subscribe to a publication
Eventually, some convert to paid
They churn from one publication and discover another via the recommendation engine
They subscribe to that publication
Repeat
The model treats every subscription event equally. A subscription to Publication A has the same positive signal weight as a subscription to Publication B. (Note: There may be factors that favor publications with bestseller status in terms of author feed visibility that I’m not accounting for here.)
SAIE may be largely agnostic about which publications people subscribe to. As long as a person stays in the Substack ecosystem (and migrates to a revenue-producing publication), that’s a victory for Substack. The house always wins.
Strategies for Success
Volumes have been written about how to gain attention for your work on Substack. A lot of it is very useful. I’ll confine my advice to what this analysis of SAIE suggests.
The Trend is Your Friend: As you’ve seen, SAIE offers compounding benefits for authors on the platform long-term. As you participate in the network, you gain more subscribers and followers. An increase in attention leads to higher responses to your content, this leads to additional subscriptions and the flywheel continues. SAIE helps increase visibility and subscriber momentum toward your publication. Account longevity and signal density (generated by your actions) counts.
Consistency is Essential: Being visible on the platform provides SAIE with more signals about you, your interests and the topics you write about. Think of posting Notes, comments, replies and articles as depositing into the SAIE signal bank. The more signal you provide, the better the engine can work for you.
Quality Over Quantity: It can be tempting to try to juice your subscriber count by ‘hacking SAIE’. If your goal is long-term viability, a loyal readership and sustained revenue, that might not be the way to go.
Letting SAIE help you find your tribe and growing your visibility among people who want to hear from you is always the superior option.
Infographic Resource
Thanks for reading. I hope you found this essay helpful. If you’re interested in receiving additional first principles-based, no-hype AI analysis, education, tools and strategy, please subscribe.
This newsletter is part of the Doing AI Efficiently Operating System, built on five operational layers: Grasp, Discern, Ward, Execute, and Honor. This essay is part of the Grasp layer, which is focused on helping you understand how AI works from first principles, including tokens, context windows, model architecture, how LLM content generation happens, and more.







