Today, The Information reported that a number of large technology companies, including Palantir, Booz Allen and Nvidia, are reconsidering their use of frontier AI models from OpenAI and Anthropic. The reason? Concerns over data security and privacy and whether AI labs are “misusing their intellectual property.”
People have always had privacy and IP concerns about using Claude, Codex, Gemini and other AI models. Do you own your work product if Claude helped create it? What about your proprietary data and methods? Would data be retained or used to train models, or even develop competing products and services?
AI labs’ terms of service state that you own your IP. Model training and data retention is a little more complicated. Certain ChatGPT plan types turn on data retention and model training by default. Business accounts generally have strict data retention and training prohibitions.
These assurances helped build trust that organizations, teams and individuals could use frontier models without worrying about these issues. Recent developments are starting to change people’s minds.
OpenAI Lawsuits Open Up User Chat Logs to Retention and Scrutiny
OpenAI is currently facing multiple lawsuits over how the company handles user data. Importantly, some of this litigation is requiring the company to preserve user chat logs, regardless of its normal data retention policy.
This means that user’s redacted chat logs, which includes highly sensitive information, such as health details, company strategic planning and other topics can be requested by litigants and defendants. Some of this de-identified information may be referenced in court filings, transcripts and other materials. Another consequence is that content “typed into ChatGPT could surface in unrelated litigation [such as divorce] cases, employment disputes and contract lawsuits.”
The Mathematics Controversy That is Raising Eyebrows
Another driving force in data privacy concerns is the recent controversy over a mathematical problem (Navier–Stokes) co-solved by Tristan Buckmaster. During the year-long process of working on the solution, they used Claude, Codex, GPT-5.6 Sol and Astra. The work is an example of how an LLM can be used to speed up progress in mathematics significantly.
OpenAI received a rumor that someone had solved the problem. 88 hours later, OpenAI had solved it as well.
Buckmaster was concerned OpenAI had used his prompts to solve the problem and asked: “whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project.” He was told “the model did not look up user data.”
Later in a statement, OpenAI said that it “cannot rule out that de-identified data derived from the usage of our products helped improved our models.”
The Data Sovereignty Imperative
It is still unclear what happened here, but these events have increased concerns about sharing proprietary data with Open AI and other labs. Now organizations are rethinking their willingness to use these models, as indicated by The Information report.
Another driver of this trend is model costs. While AI inference is still heavily subsidized, the latest models are much more expensive. And, organizations are leaning away from using them due to cost concerns. This is partly because less expensive models are more than capable of assisting with many tasks.
These are the reasons why I believe we are going to start hearing a lot more about companies pushing to secure private inference. These can be models run in data centers on their own rented/owned machines, cloud instances with strong zero data retention policies, or even spending on ‘AI computers’ that can be run on-premises.
Traditionally organizations, especially in sensitive industries like healthcare, financial services and security, have had tight data access and sharing policies. It’s highly unusual for a range of companies to voluntarily share their most sensitive data with third parties.
As access to powerful AI inference becomes more democratized, we’ll see a return to form with data/IP sovereignty, protection and access being emphasized.
This newsletter is part of the Doing AI Efficiently Operating System, built on five operational layers: Grasp, Discern, Ward, Execute, and Honor. This essay is part of the Ward layer, which gives you knowledge, tools and software to guard against AI privacy and security risks.
Additional resources in the Ward layer include the AI Security Action Pack, which features 15 in-depth guides on AI security.



