update · TowCue Editorial Team

OpenAI adds business-value analytics for ChatGPT Work and Codex: what teams should measure beyond AI usage

OpenAI's new Admin Console analytics connect ChatGPT Work and Codex usage with tasks, spend and engineering outcomes. TowCue explains what teams should measure beyond seats, prompts and token volume.

Original editorial contentSources verifiedLast reviewed: 2026-09-17

Quick answer

OpenAI published new guidance on September 16, 2026 showing how analytics in the ChatGPT Admin Console can connect ChatGPT Work and Codex activity with usage, spend, task categories and engineering outcomes. The Usage view covers active users, credits and tokens; Insights groups sampled messages into use cases and tasks; and the Outcomes view can show Codex contributions to merged commits and lines of code alongside code-review activity.

The TowCue takeaway is simple: AI adoption is not the same thing as AI value. Seats, prompts, credits and tokens tell you whether people are using a system. They do not tell you whether work got faster, better or cheaper. Teams should use activity data to choose a workflow, then pair it with a real operational outcome such as review time, defects, rework, cycle time or sales preparation time.

For related TowCue context, see the ChatGPT tool profile, Codex tool profile, and AI agent best practices.

What OpenAI added to the measurement workflow

OpenAI's September 16 product guidance describes several layers of analytics for ChatGPT Work and Codex administrators.

The Usage view brings together active users, credits and token usage. Admins can filter by group or user to see where adoption is concentrated and where a team may need training or a clearer starting workflow.

The Insights area goes beyond raw consumption. A task classifier groups a sample of messages into use cases and tasks, such as software engineering, feature development, code maintenance, sales account research and planning. OpenAI also exposes breakdowns by model, reasoning and speed, plus plugin and skill activity.

For engineering, the Outcomes view tracks Codex contributions to merged commits and lines of code alongside code-review activity, with filters including group, user and repository.

That is a meaningful shift from asking only, "How much AI did we use?"

Usage is a diagnostic, not an ROI metric

A high number of tokens can mean a workflow is valuable. It can also mean the workflow is inefficient, poorly scoped or repeatedly correcting itself.

Likewise, low usage does not automatically mean a tool failed. It may mean users do not have access, do not know which task to start with, or have not seen a workflow that saves enough time to change their habits.

TowCue would therefore treat usage data as a diagnostic layer.

It can help answer:

  • Which teams are actually adopting AI?
  • Which tasks consume the most credits?
  • Which models and reasoning settings are used for those tasks?
  • Which plugins and skills are becoming part of repeatable workflows?
  • Where should training or workflow redesign start?

But the next question must be about an outcome.

Codex shows the difference between contribution and impact

OpenAI's Codex Outcomes view can show the share of merged commits and lines of code with Codex contributions. That is useful evidence of adoption inside the software-delivery process.

It is not enough to conclude that engineering productivity improved.

OpenAI explicitly suggests comparing those trends with review time, defects and rework. TowCue agrees with that framing.

If Codex contributes to more merged code while review time falls and defects stay flat or improve, the signal is stronger. If AI-assisted code volume rises while rework and incidents rise faster, more output may simply be creating more downstream work.

The useful unit is therefore not "lines written by AI." It is something closer to accepted change that survives review and production.

Measure the workflow, not the model

The most actionable part of OpenAI's guidance is its recommendation to choose a common task tied to a business priority, agree on a baseline and an outcome with the business owner, and set a date to review progress.

That is a better evaluation pattern than comparing models in isolation.

For example, a team testing AI-assisted account research could measure:

  1. preparation time before AI;
  2. preparation time with the workflow;
  3. factual corrections required by the reviewer;
  4. whether useful sources were found consistently;
  5. whether the saved time changes sales capacity or meeting quality.

An engineering team could use the same structure with cycle time, review time, rework, escaped defects and deployment frequency.

The model matters, but the business case belongs to the workflow.

Cost visibility needs context

OpenAI's Help Center says eligible Work and Codex usage can be reviewed by surface, model, reasoning and speed, with credit and token histories available in supported workspaces. Individual Work or Codex chats can also show estimated usage details.

There is an important limitation: OpenAI warns that chat-level usage is an estimate for the main conversation and may not include every billable action. Sub-agents, separate tasks, some tools and background activity may not appear in that chat total.

That means teams should not build a precise internal chargeback system from a single chat indicator.

Use workspace billing records for spend governance, and use chat-level analytics to understand why a workflow may be expensive.

Who should pay attention

This update matters most to enterprise AI owners, engineering leaders, finance partners and operations teams that have moved beyond a small pilot.

If your organization is still testing a handful of prompts, a complex measurement program is premature. First prove that one workflow is useful.

If hundreds of people are using ChatGPT Work or Codex, however, seat counts and monthly spend are no longer enough. You need to know which work is being accelerated, where quality is changing and whether expensive configurations are justified by the task.

A low-risk measurement loop

TowCue recommends starting with one workflow rather than trying to calculate a company-wide "AI ROI" number.

Choose a task with a visible baseline. Define one speed metric, one quality metric and one cost metric. Observe the workflow for a fixed period. Then decide whether to expand it, redesign it or stop it.

A practical loop is:

usage → task → baseline → outcome → review → decision

For a coding workflow, that might become:

Codex adoption → pull-request work → review time + defects + rework → monthly review → expand or adjust

This keeps measurement close to a decision somebody can actually make.

TowCue take: stop treating activity as proof of value

AI dashboards are becoming more sophisticated, but dashboards can still create false confidence.

The important change in OpenAI's new analytics story is not another chart. It is the attempt to connect who is using AI, what task they are doing, what it costs, and what happened afterward.

Teams should resist two easy mistakes: celebrating high usage as success, and treating low usage as failure.

Instead, use usage to locate the workflow. Use operational metrics to judge the workflow. Use cost data to decide whether the quality and time savings justify the configuration.

The management question should move from:

"How much AI did our company use this month?"

to:

"Which workflows measurably improved, what evidence supports that conclusion, and where should we invest next?"

That is a much more durable way to manage AI adoption.

Sources

Research sources

Turn this intelligence into a reusable Cue

Related decision guides