OpenAI DevDay 2025: AgentKit Marks the Platform Shift From Chatbots to Production Agents

At DevDay 2025 in San Francisco, OpenAI unveiled AgentKit, Codex GA, and cheaper realtime and image models — signaling a deliberate pivot from conversational demos to deployable agent infrastructure.

8 min read

When Sam Altman took the stage at Fort Mason in San Francisco on October 6, 2025, the framing was unmistakable: OpenAI is no longer selling a chatbot. It is selling an agent platform. DevDay 2025, the company's annual developer conference, arrived with a product stack designed to carry AI from prototype conversations into production workflows — and the centerpiece was AgentKit, a bundled set of tools that treats autonomous software as the default unit of value rather than a single-turn reply.

The timing matters. ChatGPT now reports 800 million weekly active users, a scale that turns every platform decision into an industry event. Yet Altman's keynote spent comparatively little time celebrating that metric. Instead, he walked through a narrative arc that began with GPT models, moved through assistants and plugins, and landed on agents that plan, call tools, persist state, and coordinate with one another. AgentKit is OpenAI's attempt to own that entire stack — visually, administratively, and commercially.

AgentKit: From Canvas to Connector Registry

AgentKit is not a single product name on a pricing page. It is a family of components aimed at different personas inside an engineering organization — the builder who designs flows, the admin who governs integrations, the product team that embeds chat, and the reliability engineer who measures whether agents actually work.

The most visible piece is Agent Builder, a visual canvas for composing multi-agent workflows. Where earlier OpenAI tooling often pushed developers toward imperative code or JSON configuration, Agent Builder presents agents as nodes on a graph: triggers, planners, tool-using workers, human approval steps, and output formatters. The interface echoes workflow automation products, but the underlying primitives are LLM-native — branching on model reasoning, delegating subtasks to specialized agents, and passing structured context between steps rather than merely piping strings.

Alongside the canvas sits the Connector Registry, which OpenAI positioned as central administrative infrastructure for data and tool connections across both ChatGPT and the API. The registry consolidates what had been scattered integration setup into one place where organizations can approve, audit, and rotate credentials. Out of the gate, OpenAI highlighted first-party connectors to Dropbox, Google Drive, Microsoft SharePoint, and Microsoft Teams — a clear enterprise pitch — while also supporting third-party connections through MCP (Model Context Protocol) servers. For teams already standardizing on MCP as an interoperability layer, the registry effectively becomes a control plane: which agents may reach which systems, under what scopes, and with what logging.

ChatKit, the third major AgentKit pillar, addresses a problem every product team recognizes once an agent works in a demo: embedding it without rebuilding UI from scratch. ChatKit offers embeddable chat interfaces that inherit OpenAI's interaction patterns while remaining customizable enough for branded experiences. The goal is to shorten the path from a functioning agent graph in Agent Builder to a customer-facing surface on a website or internal portal.

Expanded Evals and the Production Mindset

Perhaps the least flashy but most strategically important AgentKit announcement was expanded Evals for agents. OpenAI has offered evaluation tooling before, but agent evals introduce dimensions that single-turn benchmarks miss: Did the agent choose the correct tool? Did it recover from a failed API call? Did it hallucinate a policy when accessing a document repository? Did latency stay within an SLA when the workflow fanned out to multiple agents?

By elevating evals alongside Builder and Connectors, OpenAI is signaling that "works on my laptop" is no longer sufficient. Enterprise buyers — the same buyers targeted by SharePoint and Teams connectors — demand regression suites, red-team scenarios, and measurable quality gates before agents touch customer data. The expanded eval framework also creates a feedback loop with model improvements: failure modes discovered in production-like tests can inform fine-tuning, prompt templates, and default tool policies.

Industry observers noted that this trio — visual orchestration, governed connectivity, and rigorous measurement — mirrors how cloud vendors matured container platforms a decade ago. Kubernetes did not win because it was the first orchestrator; it won because it bundled scheduling, networking, and observability into a coherent story. AgentKit reads as OpenAI's bid for a similar coherence in the agent era.

Codex Goes GA: Slack, SDK, and the Developer Loop

DevDay was not exclusively about AgentKit. OpenAI also declared Codex generally available, expanding a coding agent story that had been in preview for months. Codex GA ships with Slack integration, allowing teams to invoke coding tasks from channels where engineering work already happens — triaging bugs, proposing patches, explaining diffs, and opening pull requests without context-switching to a separate IDE pane.

The Codex SDK matters for the same reason ChatKit matters for product teams: embedding. Organizations that want coding assistance inside proprietary developer portals, CI dashboards, or internal CLIs can integrate Codex programmatically rather than forcing developers onto OpenAI's surfaces. GA status implies stability commitments, pricing clarity, and enterprise support pathways that previews typically lack.

Together, Codex and AgentKit sketch a full-stack picture — agents that write code, agents that operate business workflows, and shared infrastructure for connecting both to corporate systems. Altman emphasized that the boundary between "developer tool" and "business automation" is dissolving as natural language becomes the primary interface to software creation and operation.

Cheaper Realtime and Image Models

Not every announcement was about orchestration. OpenAI also introduced gpt-realtime-mini and gpt-image-1-mini, positioned as cost-optimized variants of its realtime voice and image generation capabilities. The company claimed gpt-realtime-mini is roughly 70 percent cheaper than its predecessor tier, while gpt-image-1-mini is about 80 percent cheaper on image tasks.

Pricing moves at DevDay are easy to underplay next to marquee platform launches, but they shape adoption curves. Realtime voice agents — customer support, coaching, accessibility tools — live or die on marginal cost per minute. Image generation at scale — marketing variants, catalog enrichment, game assets — behaves similarly. By cutting unit economics sharply, OpenAI lowers the barrier for developers to leave experiments running in production, which in turn feeds usage data and strengthens the platform moat.

Analysts compared the mini model strategy to how hyperscalers use burstable instance types: not every workload needs the flagship model, but every workload needs a predictable bill. For startups building agent businesses on thin margins, the mini tiers may determine viability.

Fort Mason, 800 Million Users, and the Platform Bet

DevDay's venue — Fort Mason on San Francisco's northern waterfront — carried symbolic weight. After years of virtual keynotes and distributed hackathons, OpenAI returned to an in-person developer gathering in the city most associated with the AI boom. Demo stations, partner booths, and hallway conversations about MCP servers recreated the energy of pre-pandemic Apple WWDC or Google I/O, albeit with a narrower focus: agents, agents, agents.

The 800 million weekly active users figure for ChatGPT functioned as backdrop rather than headline. It establishes distribution — any platform feature OpenAI enables can reach a population larger than most countries. But distribution without developer tooling produces novelty, not infrastructure. AgentKit is the company's answer to the question of what happens after a billion people have tried chat: how do builders monetize, govern, and maintain autonomous systems on top of that attention?

Competitors are not standing still. Anthropic originated MCP, which OpenAI now embraces in the Connector Registry. Google, Microsoft, and Amazon continue to push their own agent frameworks tied to cloud marketplaces. Startups such as LangChain, CrewAI, and countless vertical SaaS players offer orchestration layers with varying degrees of lock-in. OpenAI's advantage remains model quality and consumer reach; its risk is being perceived as a model vendor that rents the stack above the API layer to partners who capture long-term customer relationships.

What Developers Should Watch Next

For teams evaluating AgentKit, the practical checklist starts with integration governance. If your organization already stores documents in SharePoint and collaborates in Teams, the Connector Registry may reduce time-to-production compared with bespoke OAuth flows for every agent. If your risk team requires human-in-the-loop approvals, Agent Builder's graph model may map cleanly onto existing BPM diagrams — or it may duplicate tooling you already pay for. There is no universal answer; there is a clearer path to pilot.

Second, treat expanded agent evals as a requirement, not a nice-to-have. Agents that access live CRM or HR systems can cause real harm with a single wrong tool call. Building eval suites before launch — mirroring how mature teams test payment flows — is now aligned with OpenAI's product direction rather than fighting it.

Third, watch Codex GA in Slack if your engineering org lives in chat. The SDK path matters if you need white-label coding assistance. Either way, coding agents and business agents will share credentials, logging, and compliance policies; planning them in silos creates security debt.

OpenAI DevDay 2025 will be remembered as the year the company named its platform bet out loud. Chatbots introduced the world to large language models. AgentKit is OpenAI's attempt to ensure the world runs on OpenAI-shaped agents — built visually, connected safely, embedded everywhere, and measured honestly. Altman's keynote did not claim victory; it claimed responsibility for the next layer of the stack. The developer audience in Fort Mason applauded, opened their laptops, and went back to the harder work of making agents that deserve that applause in production.

More in openai

Comments

Loading comments…

Across the Network