AI news · October 8, 2026

OpenAI and Ironclad turn contract workflows into computer-use training tasks

computer-useenterpriseevaluationslegal-tech

OpenAI says GPT-6 Astra scored 32% higher than GPT-5.6 Sol on Ironclad's contract tasks while taking 48% less time per attempt. The result is an evaluation claim, not a general automation guarantee.

reported score gain over GPT-5.6 Sol
32%
reported lower time per attempt
48%
Made With Models illustration for this story

OpenAI described an October 6 partnership with Ironclad that turns complex contracting work into training and evaluation tasks for computer-use models. The workflows include configuring agreements, navigating approvals, and applying reusable legal terms, which gives the model a structured environment with real business rules instead of a generic browser benchmark. OpenAI says GPT-6 Astra scored 32% higher than GPT-5.6 Sol on its evaluation and used 48% less time per attempt.

The important idea is the evaluation design. A domain-specific task set can expose failures that disappear in a simple click demo: wrong terms, missed approvals, incomplete records, or actions taken in the wrong order. The result is still an OpenAI claim, and contract work needs human review. Teams building computer-use systems should copy the method by turning their own high-risk workflows into repeatable tasks with clear pass and fail conditions.

What you can do with it

Choose one workflow with real rules and a safe test environment. Record every required field, approval, and final state, then run the model against the same cases after each change. Keep write actions disabled until the model can explain and pass the full workflow consistently.

Our take

The partnership is useful because it treats computer use as a domain evaluation problem, not only a browser-control problem. That is the right direction for serious automation. The reported scores matter less than whether the task set, failure labels, and human review process can be repeated outside Ironclad.

Source: OpenAI ↗ — Made With Models writes the brief; the reporting is theirs.