OpenAI has detailed a research collaboration with contract-management company Ironclad aimed at improving how AI agents use specialized business software. The project converts real contracting workflows into structured training and evaluation tasks, giving OpenAI a way to test whether computer-using models can follow business rules, execute multi-step processes and verify that their work satisfies the original request.
The collaboration focuses on tasks such as configuring agreements, approval flows and reusable legal terms. OpenAI says Ironclad helped define realistic workflows and success criteria rather than supplying a generic software benchmark. That distinction matters because enterprise computer use is less about clicking through an interface and more about preserving policy, permissions and process constraints across a sequence of actions.
OpenAI says GPT-6 Astra is its first frontier model trained on these Ironclad tasks. On OpenAI's own research evaluation, Astra's average score was 32% higher than GPT-5.6 Sol's while estimated time per attempt was 48% lower. Those figures are useful indicators of the direction of the work, but they remain first-party results. No independent evaluation of the Ironclad task set was found during this run, and OpenAI has not presented the collaboration as a broad benchmark proving equivalent gains across other enterprise applications.
The work illustrates a shift in agent development from general computer control toward domain-specific operational competence. In contracting, an agent may need to understand an organization's approval rules, distinguish reusable clauses from one-off terms and confirm that a completed workflow matches the user's intent. Training against those constraints could make computer-use systems more useful in professional environments where a superficially successful action can still be wrong if it violates process requirements.
OpenAI describes Ironclad as the first partner in a broader effort involving software companies with deep knowledge of specialized workflows. The announcement therefore points to a training strategy as much as a single integration: convert expert software workflows into environments where agents can practice, be measured and improve. The important caveat is that the reported performance improvement is vendor-reported and specific to OpenAI's research evaluation. External evidence will be needed to establish how well those gains transfer to production deployments and other business systems.