AI

Inside NTT DATA’s Codex Rollout to 9,000 Employees

Inside NTT DATA’s Codex Rollout to 9,000 Employees

OpenAI case study graphic for NTT DATA Group's enterprise Codex deployment

The number OpenAI wants you to remember is 30 minutes. An incident analysis on a critical system at NTT DATA Group — the kind of postmortem that previously consumed five experienced engineers for three days — was completed by Codex in half an hour, according to a case study OpenAI published this week. It is a striking compression: roughly 120 engineer-hours of specialist forensic work collapsed into the time it takes to sit through a status meeting.

But the headline benchmark is the least transferable part of the story. Incident analyses vary wildly in difficulty, and a vendor case study naturally showcases the best run. What is genuinely useful in the NTT DATA account — and worth reading closely if you are responsible for an AI rollout anywhere — is the deployment mechanics: how a Japan-based global IT services company spanning consulting, systems development and operations got an agentic coding tool into the daily work of roughly 9,000 employees, many of whom are not engineers at all.

Chat first, agents second

The sequencing is the first lesson. NTT DATA did not lead with Codex. After entering a global strategic partnership with OpenAI in May 2025, the company deployed ChatGPT Enterprise across the organization and stood up an internal OpenAI Center of Excellence to run license distribution, technical validation, use-case development and usage monitoring. In an internal survey, more than 96% of respondents said they were satisfied with ChatGPT Enterprise and more than 95% reported productivity gains.

That foundation did real work. Employees spent months building the mundane habits of working with AI — drafting, researching, summarizing — before anyone asked them to delegate whole tasks to an agent. By the time Codex arrived, the organizational question had shifted from “what is this?” to “what can I hand off?” Hiroaki Sato of NTT DATA’s AI Technology Department describes the shift bluntly: “The idea that AI can take the lead in carrying out work has had an impact similar to the arrival of ChatGPT.”

That’s a meaningful distinction because Codex is a different kind of tool. Where chat assistants respond, Codex — OpenAI’s software engineering agent, first introduced in May 2025 — independently investigates, executes, tests and revises against a defined instruction. The failure mode of handing that capability to an organization with no AI habits is predictable: it becomes a demo, not a workflow.

The proof point was operations, not feature development

Notice what the flagship use case was not: it was not “Codex wrote our new product.” It was incident analysis — bounded, evidence-rich, procedurally well-defined work with a clear deliverable. The task has a natural agentic shape: gather logs, correlate symptoms, test hypotheses, produce a report. That is why it worked as the internal proof point that, in NTT DATA’s telling, “quickly gained attention from senior leaders” and built momentum for everything that followed.

Teams looking for their own first Codex win should copy the shape, not the domain: pick work that is tedious, structured, and verifiable, where a wrong answer is caught by review rather than shipped to customers.

The rollout escaped the engineering department

The second half of the case study is about what happened when Codex spread beyond developers, under a policy NTT DATA calls “Client Zero” — the company treats itself as the first customer for anything it might later sell to clients.

Nontechnical employees now use Codex to build lightweight tools, reorganize large file sets, analyze Excel data, summarize documents, and script repetitive processes. The mundane example in the study is the telling one: extracting transportation expenses from credit card statements and moving them into travel expense forms, with Codex reading multiple files, understanding each sheet’s structure and entry rules, and verifying the completed transfer. Another shift: analytical reports that once required an engineer to prepare data in a BI tool and build a dashboard can now be produced by the analyst directly from raw data.

None of these tasks is glamorous. All of them share a property that matters for adoption math: they were exactly the automations people always wanted but could never justify engineering time for. That is where an agent that writes and runs its own code earns its seat — the backlog of small automations below the priority cutoff of every IT department on earth.

The measured effect on adoption is one of the study’s concrete data points: weekly active Codex users grew 1.4x after the company published a usage guide and ran hands-on training. Tooling alone did not spread itself; documentation and training moved the number. NTT DATA also automated internal system operations with Playwright and packaged the automations as reusable Skills — turning one team’s script into organizational infrastructure that any department can pick up, rather than a private convenience that dies when its author changes roles.

The Center of Excellence’s role in that loop is easy to miss and hard to overstate. Its job, per the case study, is to identify high-impact use cases inside individual departments, generalize them, and re-share them as best practices with appropriate safeguards attached. That is the difference between an organization with 9,000 people independently prompting an agent and an organization that compounds what its best users discover. NTT DATA’s stated lessons make the same point from the other direction: deploy broadly enough to create network effects through peer learning and word of mouth, and treat deployment as the beginning of the program — then keep improving it with usage data, surveys and employee interviews — rather than the end of it.

Governance came before scale, not after

The part most enterprises get backwards, NTT DATA appears to have done in order. Before pushing Codex beyond early teams, the Center of Excellence wrote security guidelines covering what data can be used, which systems Codex may connect to, how network traffic is managed, which sandbox mode applies, what level of automation is appropriate, and where human review is required.

That list is a compact checklist for anyone deploying agentic tooling. Sandbox mode and system connectivity are decisions, not defaults; “where is human review required” is a question to answer per workflow, not per tool. Yuji Shono, who heads NTT DATA’s Global AI Office, frames governance as the enabler rather than the brake: “Creating a secure and well governed environment is essential for employees to use Codex with confidence,” he says, adding: “We want it to be a natural part of everyday work for every employee, including those in nontechnical roles.”

The claim is easy to nod past, but the causality matters. The reason 9,000 employees — technical and not — can be allowed to run an agent that executes code is that the guardrails were specified first. Enterprises that pilot agents informally and try to retrofit policy later tend to end up with either a lockdown that kills adoption or an incident that kills the program.

How to read a vendor case study like this one

The usual caveats apply. This is OpenAI marketing its own product through a strategic partner’s success story; there is no failure data, no cost accounting, and no baseline for how often Codex runs produce wrong or unusable output at NTT DATA. The 30-minute incident analysis is one anecdote, not a distribution. And NTT DATA is unusually well-positioned to succeed — it is an IT services firm whose employees professionally integrate technology, operating under a partnership that gives it direct access to OpenAI support.

There are also questions the study simply does not answer. Nine thousand employees is a large deployment, but the case study never says what fraction of the total workforce that represents, how usage is distributed across that population, or what the license and infrastructure spend looks like against the hours recovered. “More than 95% reported productivity gains” is a self-reported survey figure, not a measured output delta. None of that makes the account worthless — it makes it a deployment narrative rather than an ROI proof, and it should be weighed accordingly by anyone building a business case on top of it.

What survives the discount is the playbook, because it does not depend on the benchmark being typical. Sequence chat before agents so habits precede delegation. Pick a first use case that is structured and verifiable. Write the security guidelines before the rollout, not after the incident. Invest in guides and training, and measure whether they move active usage. Package working automations into shared, reusable form. Extend access beyond engineers deliberately, because the deepest backlog of automatable work sits with the people who could never write the script themselves.

NTT DATA’s own framing of the endgame is worth taking seriously as a signal of where enterprise AI procurement is heading: the company says employees now use ChatGPT to think and Codex to execute — people set direction and evaluate results while the agent moves work forward. Whether or not that division of labor becomes the standard, a major global IT services company is now running it live on 9,000 seats and packaging the lessons for its clients. The interesting competition in enterprise AI this year is not between models. It is between rollout playbooks — and this one is unusually well documented.

We may earn commission from affiliate links at no extra cost to you. Last updated: Jul 25, 2026.
Jinultimate

Editor of ZBrandCo and the person accountable for what we publish — setting our sourcing standards, fact-checking claims against primary sources, and issuing corrections promptly across AI, open source, and gaming. Reach the desk at editorial@zbrandco.com.