Microsoft chief executive Satya Nadella has published a long essay arguing that companies should stop treating frontier AI models as trustworthy black boxes and start treating them the way security teams treat powerful employees: as insider risks. The piece, titled "Models as Insider Risks in the Super Intelligence Era", went up on his X account on October 10 and was quickly picked up by TechCrunch and The Verge, mostly for one line. "We must assume a model is compromised and contain it from the start," Nadella wrote. "Think of it like an emergency brake."

His starting point is a gap in understanding. With conventional software, engineers could trace a behavior back to a specific code path. With today's frontier systems, Nadella writes, nobody can attribute an output to particular training data or weights, yet these systems are being given sensitive data and the ability to take consequential actions for the organizations that run them. A model provider's assurances, he argues, do not move responsibility away from the deployer.

The core proposal is to "separate the supply of intelligence from the authority over it." In practice that means keeping the model apart from the harness that orchestrates its work and from the action space that defines what it may do, and placing controls and safeguards outside the model. Nadella ties this to an old information-security rule: a program should never be able to bypass or tamper with the mechanism that enforces its own permissions.

He then lists seven design principles. No single model should be the sole dependency for an important outcome or check its own work. Every meaningful model action should leave tamper-proof, human-readable evidence, so an outcome can be reconstructed without asking the model to vouch for itself. Whole systems should be tested continuously, including attacks and failures. Organizations should be able to decide independently what a model can reach and do, and validation should not be performed by the intelligence being validated. An authorized person should always be able to pause or stop a model mid-task, and Nadella says more capable models will need containment technology the industry should standardize. Finally, failures should be disclosed promptly to those affected, with details of which controls broke shared across the industry.

He also calls chain-of-thought transparency "non-negotiable" while conceding it is not enough on its own, since models' stated reasoning is not reliably faithful. Using models to test each other is useful, he writes, but can leave "an opaque model inside an opaque orchestration layer, watched by another opaque model."

The timing is notable. The essay arrived a day after Anthropic published a report on unintended actions by its Claude models during evaluations and internal use, and TechCrunch placed it in a run of incidents in which AI companies appeared to lose control of their systems. The Verge observed that most of Nadella's recommendations, such as incident disclosure, independent audits and containment, echo what others in the industry have proposed, with the explicit shutdown requirement going somewhat further.

The essay is a statement of principles, not a product or policy announcement. It does not name any Microsoft product, commit the company to a timeline, or explain how the principles map onto Microsoft's own AI services. Nadella also uses the term "Super Intelligence" throughout for today's AI systems, a framing both outlets noted. Whether the ideas turn into standards or contractual commitments from Microsoft and its peers remains to be seen.

References