The race to build increasingly powerful artificial intelligence models is entering a new phase, with Elon Musk calling for rival AI companies to review each other’s most advanced systems before they are released to the public.
Musk’s proposal would place competing AI laboratories in a position that may seem unusual for the technology industry: reviewing one another’s models for potential safety and security problems before those models reach users.
The idea comes as companies including OpenAI, Google DeepMind, Anthropic and Musk’s xAI compete to develop increasingly capable AI systems.
Musk told The Economist that leading AI companies should meet regularly to discuss safety and security issues and proposed peer review of frontier models before release. He also argued that governments should become involved if companies fail to address serious risks themselves.
The proposal highlights a growing question for the AI industry: Can companies developing the world’s most powerful AI systems effectively regulate one another, or will independent government oversight eventually be necessary?
What Elon Musk is proposing
Musk’s idea is relatively straightforward.
When an AI company develops a frontier model that is substantially more capable than existing systems, rival laboratories could be given an opportunity to evaluate it before public release.
The purpose would not necessarily be to prevent competitors from launching products. Instead, the review process could identify serious safety or security problems that the developing company may have missed.
Musk has also suggested that leading AI companies should regularly communicate about emerging risks rather than treating safety entirely as an internal matter.
The proposal effectively creates a form of industry-level peer review.
That is significant because the AI industry currently relies heavily on internal testing, voluntary safety commitments and company-specific evaluation procedures.
Why AI peer review matters
Traditional software can often be tested against clearly defined requirements.
Frontier AI systems are more complicated.
A model can generate code, interact with external systems, reason through complex tasks and operate with varying degrees of autonomy. As capabilities increase, developers may not always be able to predict every way a system could behave when deployed at scale.
Independent testing can therefore provide another layer of scrutiny.
A rival AI company may also have expertise that the original developer does not. Different laboratories use different training techniques, evaluation methods and safety approaches.
That diversity could help expose weaknesses.
There is already movement toward broader external oversight. Google DeepMind CEO Demis Hassabis has proposed an independent standards body that could test frontier models before release, while other AI leaders have also called for stronger oversight mechanisms.
But should AI companies be trusted to police each other?
This is where Musk’s proposal becomes more complicated.
OpenAI, Anthropic, Google DeepMind, xAI and other companies are not neutral observers.
They are commercial competitors.
Each company has significant financial and strategic incentives to release increasingly capable products quickly. Delaying a model because a competitor identifies a safety concern could therefore have commercial consequences.
There is also the question of confidentiality.
Frontier AI models represent enormous investments in computing infrastructure, research and engineering. Companies may be reluctant to give competitors access to their systems before launch.
A workable peer-review system would therefore need strict rules around access to models, confidential information, intellectual property and disclosure of vulnerabilities.
Peer review would not replace independent regulation
Musk’s proposal is particularly interesting because it does not necessarily reject government oversight.
Instead, his position puts industry cooperation first, with government intervention becoming more important when companies fail to respond to serious risks.
That creates a potential middle ground between two extremes.
One approach is almost entirely voluntary self-regulation.
The other is extensive government control over AI development.
A third model could combine the two: companies conduct technical reviews while independent regulators establish minimum requirements, investigate serious incidents and enforce rules when voluntary measures fail.
That distinction matters because an AI company reviewing its competitor is still not the same as an independent regulator.
The timing is significant
Musk’s comments came during growing concern about the behaviour and safety of frontier AI systems.
The debate has intensified following reports involving AI systems interacting with external computer systems in ways that raised questions about the effectiveness of existing safeguards.
More recently, Anthropic CEO Dario Amodei called for the AI industry to slow the pace of capability development and proposed independent evaluators with access to AI companies’ internal safety processes. Musk and OpenAI CEO Sam Altman have expressed support for parts of Amodei’s approach.
OpenAI has also called for mandatory national AI safety requirements, including independent assessments, cybersecurity measures and incident reporting for advanced AI systems.
That suggests the industry debate is moving beyond the question of whether AI should be tested.
The bigger question is now who should test it, when should testing happen and what happens when a dangerous capability is discovered?
What would an AI peer-review system actually examine?
A serious review system would need to go beyond asking whether an AI chatbot produces incorrect answers.
Potential evaluations could include:
- Cybersecurity capabilities
- Autonomous computer use
- Ability to manipulate or deceive users
- Privacy and sensitive-data risks
- Dangerous biological or chemical information
- Ability to circumvent safety controls
- Model behaviour when given conflicting instructions
- Resistance to prompt injection and other attacks
- Reliability when operating autonomously
- Ability to replicate or improve parts of its own operation
- Potential misuse by criminals or other malicious actors
The exact tests would depend on the capabilities of the model.
A simple chatbot and an AI agent capable of independently operating software should not necessarily face identical evaluations.
The biggest challenge: agreeing on what is dangerous
Another problem is that AI companies may not agree on what constitutes an unacceptable risk.
One laboratory might consider a particular capability manageable with additional safeguards.
Another might consider the same capability too dangerous for public deployment.
That means a peer-review system needs objective thresholds.
For example, reviewers could identify specific capabilities that trigger additional testing or require a delay before deployment.
Without measurable criteria, peer review could become little more than another voluntary discussion between companies.
What does this mean for AI users in Kenya?
The debate may appear to be taking place mainly in Silicon Valley, but its consequences will extend far beyond the United States.
Kenyan businesses, universities, developers, government agencies and consumers are increasingly using AI tools for customer service, software development, education, research, marketing and financial services.
As more powerful AI systems become available globally, Kenyan users will inherit many of the same risks.
For businesses, a major concern is not simply whether an AI model produces an incorrect answer. It is whether an AI agent connected to company systems can take an incorrect or harmful action.
For example, an AI system connected to customer records, payment systems or internal software could create a much larger problem than a conventional chatbot making a factual mistake.
This makes independent testing and clear accountability increasingly important.
Kenya will also need to think about AI evaluation
Kenya does not need to reproduce Silicon Valley’s regulatory model exactly.
However, the country will increasingly need mechanisms for evaluating AI systems used in sensitive areas.
That could include financial services, healthcare, education, public services, telecommunications and cybersecurity.
A practical approach could involve cooperation between regulators, universities, technology companies and independent researchers.
Rather than waiting for serious incidents before establishing testing frameworks, Kenya could develop expertise in AI evaluation as adoption grows.
That would also create an opportunity for African researchers and technology institutions to participate in the global AI safety conversation rather than simply importing standards developed elsewhere.
Could rival AI companies really cooperate?
That may ultimately be the hardest part of Musk’s proposal.
AI companies have strong reasons to compete.
The industry is moving extremely quickly, and a company that delays the release of a powerful model could potentially give competitors an advantage.
At the same time, the companies share some risks.
A major AI safety incident could damage public confidence across the entire industry, not just the company responsible.
That creates an unusual situation where competitors have an incentive to cooperate on certain safety issues while continuing to compete aggressively on products and capabilities.
The aviation industry offers a useful analogy: airlines compete commercially, but they operate within shared safety standards because a major safety failure can affect confidence in the wider sector.
AI may eventually require a similar distinction between competition in products and cooperation on safety.
The case for independent reviewers
While industry peer review could be useful, independent reviewers could provide an additional layer of accountability.
An independent evaluator would ideally have sufficient technical expertise and access to conduct meaningful testing without being financially dependent on the company whose model is being assessed.
That could reduce the risk of conflicts of interest.
Recent proposals from AI industry leaders increasingly point in this direction. Anthropic’s Dario Amodei has called for independent evaluators to have significant access to AI companies’ safety processes, while OpenAI has backed mandatory independent assessments as part of national AI safety requirements.
The AI industry is moving toward a new safety debate
Musk’s proposal is important not because peer review is guaranteed to become the industry’s standard, but because it illustrates how quickly the AI safety conversation is changing.
A few years ago, much of the public debate focused on whether AI regulation was necessary at all.
Today, leading AI companies and executives are increasingly discussing independent evaluations, standards bodies, incident reporting and stronger safety requirements.
The disagreement is increasingly about how much oversight is appropriate and who should have the authority to enforce it.
That is a much more consequential debate.
What happens next?
For Musk’s idea to become practical, the AI industry would need to agree on several difficult issues.
Who qualifies as a frontier AI laboratory?
How much access should competitors receive?
How long should a review take?
What happens when reviewers identify a serious vulnerability?
Who decides whether a model should be delayed?
How are confidential technologies protected?
And, most importantly, what happens when an AI company refuses to address a serious risk?
Those questions could determine whether peer review becomes a meaningful safety mechanism or simply another voluntary industry initiative.
Conclusion
Elon Musk’s proposal to have rival AI companies peer-review one another’s most advanced models represents an unusual attempt to bring cooperation into an industry built around intense competition.
The idea has potential because AI laboratories possess the technical expertise needed to test increasingly sophisticated systems.
But peer review alone is unlikely to solve the industry’s governance challenge.
The stronger model may ultimately involve multiple layers of protection: internal testing, independent evaluation, cooperation between AI companies and enforceable government standards.
For countries such as Kenya, the lesson is equally important. As AI adoption accelerates, the conversation should not focus only on how quickly businesses and consumers can adopt new AI tools.
It should also ask whether those systems have been properly tested, who is responsible when they fail and whether users have meaningful protection when AI causes harm.
The next stage of the AI race may therefore be about more than building smarter models.
It may also be about building better systems for deciding when those models are safe enough to deploy.
Frequently Asked Questions
What did Elon Musk propose about AI peer review?
Elon Musk proposed that leading AI companies should review one another’s most advanced frontier AI models before public release, with companies also meeting regularly to discuss emerging safety and security concerns.
Why does Elon Musk want rival AI companies to review models?
The idea is to give AI developers an additional layer of technical scrutiny and identify potential safety or security problems before powerful models reach the public.
Would AI peer review replace government regulation?
Not necessarily. Musk has argued that government intervention should become necessary when companies fail to address serious risks. Other proposals in the industry favour combining independent evaluation with mandatory regulation.
What is a frontier AI model?
A frontier AI model generally refers to one of the most capable AI systems available at a particular time, particularly models that introduce significantly greater capabilities or autonomy.
Why is AI safety important for Kenya?
Kenyan businesses and consumers are increasingly using AI for areas including business operations, education, software development, financial services and customer support. As AI becomes more capable, safety, privacy, cybersecurity and accountability become increasingly important.
Could AI companies refuse to participate in peer review?
Yes. That is one of the major weaknesses of a voluntary system. A formal framework would need clear incentives or regulatory requirements to ensure that companies participate and act on serious findings.
What is the alternative to AI companies reviewing one another?
Possible alternatives include independent AI safety organisations, government regulators, academic testing laboratories or a combination of industry and independent oversight.

