OpenAI just made two moves at once that are worth paying attention to: a significantly upgraded model and a consolidated app that folds several separate tools into one. Neither change is cosmetic.
What GPT-5.6 Actually Means
Model versioning can feel like marketing. But the jump to GPT-5.6 is a real performance gap, not a point-release polish. The short version: you’re getting near-frontier intelligence at a fraction of the cost and time that previous top-tier models required.
To put that in concrete terms — tasks that previously required several minutes and several dollars to run at high quality are now completing in roughly a minute and at well under a dollar. That’s not a small shift. It changes which workflows are actually economical to automate.
On standard capability benchmarks covering hard reasoning questions, browser-based tasks, and economically useful work, GPT-5.6 competes directly with the best available models from other labs. In software engineering and agentic terminal tasks specifically, it currently leads the benchmarks most practitioners trust.
The model comes in multiple tiers — a lighter version for fast, cheap tasks and a heavier “Soul Ultra” tier for complex, multi-step work. The lighter end is good enough for most daily use. Reserve the top tier for things like building a multi-page interactive app from a single prompt or running a deep analysis across your entire document library.
What It’s Genuinely Good At Right Now
- Generative coding: Full interactive web apps — with working physics, camera controls, and styled visuals — from a single prompt
- Hard reasoning: Tying or beating frontier models on graduate-level question-answering benchmarks
- Agentic tasks: Running terminal commands, browsing, and operating tools with less hand-holding than before
The Super App: One Interface Instead of Three
OpenAI has been running separate products for chat, coding, and browsing. That’s now collapsing into a single ChatGPT app with two modes: Work and Codex.
Think of the distinction this way:
- Codex mode is for developers actively writing and debugging code. You see branch selectors, a terminal panel, and code review tooling in the sidebar.
- Work mode is for everything else — drafting, research, managing tasks, coordinating with your tools. It looks cleaner, connects to external apps like Gmail, Slack, and Google Drive, and decides on its own whether your request needs a simple answer, a web search, or a mini software build.
The built-in browser (previously a standalone product) is now just part of the app’s sidebar. Same for hosted site publishing — you can build something, click a button, and have it live at a ChatGPT-hosted URL without touching a server.
The Personal Assistant Layer
Work mode’s most interesting capability is what happens when you connect your actual accounts and let it survey your situation. Point it at your inbox, calendar, Slack threads, and recent documents, and ask it to identify where you’re losing time or dropping balls. It produces a prioritized list of what needs attention, what it can handle autonomously, and what’s waiting on someone else.
You can then schedule that kind of sweep to run daily. The practical result is something close to a chief-of-staff briefing every morning — what’s urgent, what’s overdue, what conflicts exist in your calendar — without you having to pull it together manually.
This only works well if you actually connect your tools. A half-connected setup gives you half the value.
GPT Live: Real Conversation, Not Just Voice Commands
The updated voice mode is worth a separate mention. Previous voice interfaces were essentially text prompts you spoke aloud. The new version handles natural interruption — you can cut in, it can cut in, and the conversation flows closer to how humans actually talk.
The translation use case is the most immediately practical: set your phone in the middle of a table, tell it which two languages to translate between, and it handles the back-and-forth in real time. No app-switching, no typing. If you work with multilingual clients or travel frequently, that’s worth testing today.
To enable it: Settings → Voice → set to Live, then choose your intelligence tier based on whether you want speed or depth.
The Competitive Picture
OpenAI isn’t alone. xAI’s Grok 4.5 is performing competitively on the same benchmarks — strong on software engineering and terminal tasks, available for local installation via a simple terminal command. Meta has also been active. The field is tighter than headlines usually suggest, which is actually good news: competition is pushing all the top providers to ship faster and price more aggressively.
GPT-5.6 leads on the benchmarks that matter most right now, but the gap to the next-best option is measured in percentage points, not leagues.
What to Actually Do This Week
If you’re already a ChatGPT user, update the app and spend 20 minutes in Work mode with your real accounts connected. Ask it to audit your last 30 days and tell you where to focus. That single exercise will tell you more about what the new setup can do for your specific situation than any benchmark will.