As AI models become more powerful, the potential for them to be exploited increases, and so does the need for safety guardrails that can prevent such exploits from occurring. AI companies currently have to walk a delicate tightrope between respecting the privacy of their business customers while monitoring usage for problems.
Sensing an opportunity to beat rival Anthropic, OpenAI just announced a privacy-centric safety approach to monitoring abuse. The company is previewing a new service called “Private Safety Processing” for some customers. This is an automated system that monitors for potential abuse without retaining any customer data.
This system clearly violates Anthropic’s recently announced data retention policy. The policy, which has made some customers uncomfortable, allows AI Labs to store user data (all sessions and conversations therein) for 30 days for “covered models.” These models include all Mythos-class models and “future models with similar features,” the company says.
The policy, announced in July, is designed for safety purposes and allows laboratories to screen for and analyze potential misconduct. However, this is a major concern for some companies that handle large amounts of sensitive data and do not want their data stored (or examined) by an AI lab.
OpenAI, like most other AI companies, already offers a relative level of privacy to its customers by adhering to a policy known as zero data retention. ZDR uses agents within the OpenAI API to monitor fraud on a session-by-session basis. In this way, customer data is not retained by the company, but the company can scan for fraudulent activity without the need for human intervention. It’s worth noting that Anthropic is also largely ZDR compliant, with the exception of “covered models” like Fable.
According to OpenAI, Private Safety Processing is a new technology that expands the scope of ZDR. This is described as a form of long-term safety monitoring that evaluates the input and output of multiple conversations rather than just one. Again, monitoring is performed by the agent, which, when triggered, captures the interaction and analyzes the entire session for signs of potential abuse.
This new technology helps OpenAI detect malicious uses of AI that occur across multiple sessions, a spokesperson told TechCrunch. A malicious attacker (for example, someone creating malware for a cyber attack) could scatter requests to avoid detection. Private Safety Processing can analyze these multiple conversations for signs of abuse without requiring a human to see the user’s conversations.
If the system is triggered, it could send a “narrow signal” to OpenAI that alerts it to certain types of activity, the company said. Based on that signal, OpenAI can determine whether enforcement is necessary. In that case, OpenAI will contact the customer to ask for more information or assist with the issue, and customers can choose to share their data with OpenAI at their discretion, the spokesperson said.
In contrast, Anthropic notes that human review of customer data may occur, but only “through a controlled access path” involving “a small number of authorized reviewers.” The company says all such review sessions are “recorded in tamper-proof logs and cannot be suppressed or modified by reviewers.”
The competition between OpenAI and Anthropic is currently tense, with both companies looking for opportunities to gain an advantage over the other. OpenAI’s second quarter growth was slower than Anthropic’s, according to recent reports. Anthropic’s annual revenue run rate is now reportedly $65 billion. Humanity Investors say it could IPO for $2 trillion, and OpenAI is also working toward an IPO.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
