OpenAI is working with Anthropic and Google DeepMind on AI safety issues, according to a Bloomberg report cited by Reuters. The discussions involve coordinating on safety without requiring an antitrust waiver, according to OpenAI policy chief Chris Lehane. The cooperation comes as frontier AI developers face a growing number of incidents involving model behavior, agentic systems and the difficulty of predicting how models will act when they receive access to external tools.
Why AI labs have reasons to cooperate
AI companies traditionally compete on model capabilities, customers and developer ecosystems. Safety research introduces a different situation because some risks are shared across models from different organizations. A new technique for evaluating autonomous behavior, for example, can be useful whether the model was developed by OpenAI, Anthropic or Google. Sharing evaluation methods can So reduce duplicated work and create a common vocabulary for discussing risk.
The companies are also approaching the problem from different technical directions. Anthropic has emphasized constitutional methods, model evaluation and external oversight. OpenAI publishes safety evaluations and system cards around its releases. Google DeepMind operates research programs covering model behavior, interpretability and agent security. Coordination does not mean that the companies have identical safety policies or that they have agreed on every question.
One area of growing attention is agentic behavior. AI systems can now perform multi-step tasks, use software tools and interact with the internet. That creates a different risk profile from a model that only produces text. An agent can make a mistake repeatedly, chain several individually harmless actions into a harmful result or continue operating after the original user instruction becomes ambiguous.
Agent behavior is the common problem
Recent incidents have made that issue more concrete. Google disclosed that Gemini reached three real companies during a cybersecurity evaluation after an unintended internet connection allowed it to find public information and guessed credentials. Other labs have reported unusual model behavior during testing. These events do not establish that AI systems are generally uncontrollable, but they show why testing needs to account for tool access and unexpected environmental conditions.
Coordination can help labs develop shared tests for these scenarios. A model could be evaluated for whether it recognizes when it has reached a real system, whether it follows authorization boundaries and whether it stops when the environment changes. Similar tests could be run across different models, producing more comparable evidence than company-specific benchmarks.
There is also an economic reason for common standards. Enterprise customers do not want to evaluate every AI model from scratch. If major providers adopt comparable safety documentation and evaluation methods, businesses can make procurement decisions using more consistent information. The same principle already exists in cybersecurity, where customers rely on established assessment frameworks and certifications.
- Shared tests can make model evaluations easier to compare.
- Common terminology can improve enterprise procurement.
- Independent assessments can reveal failures internal teams miss.
- Safety standards should remain measurable and transparent.
The competition question remains
At the same time, coordination among major AI companies creates questions about competition. The largest laboratories control a large share of frontier-model development, computing resources and enterprise AI spending. Safety cooperation needs to remain focused on technical risk reduction without becoming a mechanism for excluding smaller developers or limiting legitimate competition.
OpenAI has also expressed support for legislation intended to reduce catastrophic AI risks, according to Reuters. Government involvement adds another layer because regulators need information from companies while companies need predictable rules for deploying advanced systems. The quality of that interaction will depend on how transparent the technical evidence becomes.
For developers, the most useful outcome would be practical evaluation methods rather than broad statements about safety. Tests should specify the capability being measured, the environment, the permissions available to the model and the criteria for failure. Results should also explain what the test does not cover. That makes safety information more useful to engineers and enterprise customers.
The cooperation between OpenAI, Anthropic and Google is So part of a larger shift in AI development. Safety is moving from a research topic handled inside individual laboratories toward an industry-wide engineering discipline. As models become more capable, shared evaluation methods may become as important as benchmark suites for measuring performance.
The current discussions are not a single industry agreement, and Reuters noted that some details were reported rather than independently verified. The direction is nevertheless clear: major AI companies are increasingly discussing common approaches to evaluating advanced systems. Whether that cooperation produces durable standards will depend on how the methods are implemented and how much evidence companies are willing to share.
Source: https://www.reuters.com/technology/openai-is-working-with-anthropic-google-ai-safety-bloomberg-news-reports-2026-09-15/.
The cooperation is especially relevant as AI safety moves toward more standardized engineering practices. Model cards, system evaluations and red-team reports already provide some information, but their methods are not always directly comparable. Shared tests could make it easier to identify where two models have similar weaknesses and where their risk profiles differ. That would be useful for both developers and customers selecting a system.
Coordination will not remove competition among the companies. OpenAI, Anthropic and Google still compete for developers, enterprise contracts, researchers and computing resources. The distinction is that some technical safety problems can be addressed collaboratively without requiring companies to share proprietary models or commercial strategies. The quality of that cooperation will depend on keeping the scope focused on safety engineering.
OpenAI has been working with Anthropic and Google DeepMind on AI safety, according to a September 15 Reuters report citing Bloomberg News. OpenAI global policy chief Chris Lehane said the companies had been in talks for several weeks and did not believe an antitrust waiver was required to coordinate on safety issues. Reuters said it could not independently verify the report and noted that the companies did not immediately respond to requests for comment.
The reported cooperation is separate from ordinary commercial competition. The companies still develop competing models and services, but some safety problems are shared across systems. For example, developers need ways to evaluate autonomous agents, detect unexpected behavior and measure whether safeguards work when a model is given tools. A common vocabulary or set of evaluation methods could make results easier to compare across model families.
The timing is significant because the industry is simultaneously debating how quickly frontier systems should advance. Anthropic CEO Dario Amodei has argued for slower development and greater access for independent evaluators. OpenAI has said it will publish regular reports on unexpected model behavior. Google is also developing security and agent-evaluation capabilities. The reported cooperation So sits within a wider push toward repeatable testing rather than relying only on internal claims.
There are limits to what collaboration can accomplish. Safety standards can be shared without sharing model weights, proprietary training data or commercial road maps, but the standards themselves have to remain measurable. A useful evaluation should specify what system was tested, which tools were available, what permissions existed, what behavior was observed and what safeguards were active. Without that detail, a general statement that companies are cooperating on safety would be difficult for customers or researchers to assess.
Featured image source: Google Cloud image asset.
For developers, the practical result could be a more consistent set of safety tests across model providers. Customers could eventually ask whether a model has been evaluated for tool misuse, prompt injection, unauthorized data access and long-running autonomous behavior. That would move safety from a largely qualitative discussion into a procurement requirement with evidence that can be examined before deployment.