OpenAI proposes a scorecard for AI: measure useful intelligence per dollar

OpenAI's July 17, 2026 framework asks companies to measure AI by completed work, cost per successful task, dependability, and value at scale rather than token spend alone.

On July 17, 2026, OpenAI published “A scorecard for the AI age,” shifting the enterprise AI question from “how many seats and tokens are we using?” to “how much useful work is AI completing?” The article starts with the return-on-investment question that CFOs care about and proposes a measure it calls “Useful Intelligence per Dollar.”

The framework asks four questions: how much important work does AI complete, what is the full cost of each successful task, can people depend on the result, and does each dollar create more value as usage grows? OpenAI argues that token price alone is not enough because a cheaper model may require more retries, waiting time, or human correction.

The first step is to put the metric back into the system where work actually happens. A support team can measure cases resolved, an engineering team can measure code changes that pass tests and review, and a legal team can measure contracts reviewed accurately and on time. The article recommends starting with one workflow, defining what “done” means, and recording the outcome in the existing system.

The second step is to calculate the full cost of a successful task. That includes model and tool usage, but also attempts, latency, employee time, human review, and rework. OpenAI uses its own models as examples to argue that a higher-priced model may have a lower total cost if it reaches the quality bar in fewer attempts. Companies still need to test that claim against representative cases of their own.

Dependability is another central part of the scorecard. OpenAI suggests classifying outputs as “ready to use,” “needs correction,” or “needs escalation” because those categories show whether AI is actually reducing the work needed to finish a project. As systems move from drafting to cross-tool actions and exception handling, teams need explicit boundaries for accessible data, connected systems, human approvals, and permissions.

The final measure is value at scale: within the same workflow, are accepted tasks growing faster than total cost while quality holds or improves? This makes AI investment look more like a portfolio that can be validated in stages: explore whether the model can do the job, validate representative cases against a quality bar, then fund the integrations, governance, reliability, and change management required for production.

The article is an OpenAI management framework, not an independent industry standard. Its examples involving ChatGPT Work, GPT-5.6, and compute infrastructure also come from a vendor perspective. The practical move for businesses is to keep the four questions, then build a baseline from their own definition of success, approval records, cost data, and business outcomes.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.