OpenAI Launches ChatGPT Work: Long-Running Agentic Task Automation · history
Version 2
2026-07-11 08:10 UTC · 43 items
What
OpenAI launched ChatGPT for Work on July 9, 2026—a rebranding of Codex as a unified desktop super-app—alongside GPT-5.6, which comes in three tiers (Sol, Terra, Luna) [1]. The tool is designed to execute multi-step professional tasks over hours, integrating with Slack, Teams, Google Drive, SharePoint, email, and CRMs via plugins, and includes Scheduled Tasks and a Sites feature for publishing interactive web apps [3][2]. The rebranding caused immediate confusion among expert observers: both Ethan Mollick and Simon Willison said publicly they could not understand the distinction between ChatGPT for Work and Codex [1][4]. OpenAI published enterprise case studies from Australian Payments Plus and Deutsche Telekom around the launch to support adoption claims [5][6].
Why it matters
ChatGPT for Work is OpenAI's clearest attempt to shift from a reactive chat tool to a sustained autonomous agent embedded in enterprise workflows. The product identity confusion at launch—acknowledged even by close expert observers—points to a communication problem that could slow enterprise evaluation, regardless of the underlying capabilities.
Open questions
Can OpenAI resolve the product identity confusion between ChatGPT for Work and Codex before it affects enterprise sales cycles? [1][4]
How does human-in-the-loop approval work in practice for complex, multi-hour tasks—at what granularity does the tool pause for user review? [2]
Will the cloud/desktop interface fragmentation—Work conversations on web and mobile not appearing in the desktop interface—be resolved, and how does it affect enterprise deployments? [4]
How will competitive pricing pressure from Meta's Muse Spark 1.1 ($0.80/M input tokens) affect OpenAI's enterprise positioning for agentic workflows? [1]
Narrative
OpenAI launched ChatGPT for Work on July 9, 2026, as part of what The Neuron called a 'Super Thursday' that also included the release of GPT-5.6 in three tiers—Sol (most capable), Terra (balanced), and Luna (affordable) [1]. ChatGPT for Work is effectively a rebranding and expansion of Codex into a unified desktop super-app with web browsing, file editing, multi-agent coordination, and scheduling built in [1]. The tool is designed to stay engaged with professional projects for hours—directly addressing a limitation in the earlier Atlas Agent Mode, which stopped automated tasks after a few minutes [2]. It integrates with Slack, Microsoft Teams, Google Drive, SharePoint, email, and CRM systems via plugins, and includes a Scheduled Tasks feature for recurring automation and a Sites feature, in public beta, for publishing interactive web apps [3]. OpenAI requires human approval before the tool executes important actions [2].
The launch produced immediate confusion about what the product actually is. Grant Harvey at The Neuron noted that both Ethan Mollick and Simon Willison—two observers who follow the space closely—said publicly they could not understand the distinction between ChatGPT for Work and Codex [1]. Willison separately identified a concrete UX fragmentation: Work conversations started on web or mobile run in the cloud and do not appear in the desktop Work interface, while the desktop app can additionally access local files and desktop applications [4]. Willison found OpenAI's own explanation of this split insufficient. Harvey also flagged memory reliability as a structural problem, writing that until memory becomes reliable enough, the app will feel like 'dude where's my chat because somebody rustled all the papers around on your desk' [1].
OpenAI has positioned the launch with enterprise validation evidence. Australian Payments Plus, in a case study published July 7, reported that 77% of surveyed employees save more than two hours per week using ChatGPT Enterprise, and that Codex reduced a complex reconciliation investigation from four hours to thirty minutes [5]. AP+ employees created more than 300 custom GPTs and more than 1,000 Projects, suggesting broad organizational rollout rather than isolated experimentation [5]. Deutsche Telekom announced a partnership with OpenAI around the same time, with a stated goal of becoming an 'AI-native' telco across customer service, employee workflows, and network operations [6]. Both are OpenAI-published promotional pieces with no independent verification.
The competitive backdrop is shifting. The Neuron noted that Meta released Muse Spark 1.1 with a 1M-token context window and computer use at $0.80 per million input tokens, applying direct pricing pressure on OpenAI and Anthropic in the enterprise agentic space [1]. Harvey characterized the current competitive moment as 'Anthropic's game to lose,' reflecting a view that OpenAI's UX and naming confusion create an opening for rivals [1]. OpenAI reports more than 5 million weekly Codex users, with over 1 million using it for non-software tasks, and says nearly 100% of internal teams across finance and sales now rely on the product [3]—though these figures come from OpenAI itself and have not been independently verified.
Timeline
- 2025-10-01: Ars Technica tests Atlas Agent Mode, finding execution slow and automation stopping after a few minutes. [2]
- 2026-04-23: OpenAI publishes Codex Automations documentation describing scheduled tasks that run without user prompts and continue from existing conversation context. [7]
- 2026-07-07: OpenAI publishes Australian Payments Plus case study reporting 77% of surveyed employees save 2+ hours per week and Codex cut a 4-hour reconciliation task to 30 minutes. [5]
- 2026-07-09: OpenAI launches ChatGPT for Work (a rebranding of Codex as a unified desktop super-app) alongside GPT-5.6 in three tiers, with Scheduled Tasks, third-party app integrations, and a Sites feature in public beta. [3][2][1]
- 2026-07-10: Simon Willison flags cloud/desktop interface fragmentation: Work conversations on web and mobile do not appear in the desktop Work interface. [4]
- 2026-07-10: OpenAI publishes Deutsche Telekom partnership announcement, describing a goal to become an 'AI-native' telco across customer service, employee workflows, and network operations. [6]
- 2026-07-10: The Neuron reports broad expert naming confusion (Mollick and Willison both publicly confused about ChatGPT for Work vs. Codex) and notes Meta's Muse Spark 1.1 as a direct competitive entrant at $0.80/M input tokens. [1]
Perspectives
OpenAI
Presents ChatGPT for Work as a step toward autonomous AI completing entire professional workflows; cites 5M+ weekly Codex users, near-universal internal adoption, and enterprise case studies from AP+ and Deutsche Telekom as validation.
Evolution: Consistent promotional positioning; ambition has escalated from Atlas's limited scope to hours-long autonomous operation, now paired with a product rebranding that combines Codex and ChatGPT into one unified app.
Ars Technica (Kyle Orland)
Reports the launch largely on OpenAI's terms, contextualizing it against Atlas Agent Mode's prior limitations and noting the human-approval gate as a meaningful design choice.
Evolution: Earlier Atlas coverage was more skeptical; current coverage is more descriptive, treating extended duration and integrations as a genuine capability step.
Simon Willison
Flags concrete UX fragmentation between cloud and desktop Work interfaces; OpenAI's own explanation fails to resolve the confusion.
Evolution: Consistent independent critical voice; the naming confusion critique is now corroborated by Mollick and amplified by The Neuron.
Grant Harvey (The Neuron)
Enthusiastic about capability advances but critical of UX and branding execution; flags that expert observers including Mollick and Willison cannot parse the ChatGPT for Work vs. Codex distinction, and argues memory fragmentation undermines the super-app vision.
Evolution: New voice this pass; characterizes the competitive moment as 'Anthropic's game to lose' given OpenAI's communication problems.
Tensions
- OpenAI positions ChatGPT for Work as a coherent unified product; Willison and Mollick both say they cannot understand the distinction between ChatGPT for Work and Codex, and Willison's explanation of the cloud/desktop split found OpenAI's own clarification insufficient. [3][4][1]
- OpenAI promotes hours-long autonomous task execution while requiring human approval before important actions—the practical balance between autonomy and oversight in complex multi-step workflows is undefined in public documentation. [2][3]
- OpenAI's enterprise adoption figures (5M+ weekly Codex users, near-100% internal team adoption) come exclusively from OpenAI; no independent verification or third-party performance assessments have surfaced. [3][5][6]
Sources
- [1] 😼 OpenAI's Super Thursday — The Neuron (2026-07-10)
- [2] OpenAI wants its new tool to do your work for you and with you — Ars Technica AI (2026-07-09)
- [3] ChatGPT is now a partner for your most ambitious work — OpenAI Blog (2026-07-09)
- [4] Quoting OpenAI — Simon Willison (2026-07-10)
- [5] Australian Payments Plus moves faster with ChatGPT and Codex — OpenAI Blog (2026-07-07)
- [6] How Deutsche Telekom is rewiring telecommunications with AI — OpenAI Blog (2026-07-10)
- [7] Automations — OpenAI Blog (2026-04-23)