GPT-6 is out, OpenAI’s CEO is using the word AGI in press statements, and your feed is full of three.js demos. Before you upgrade your subscription or restructure your workflow around it, here’s a clear-eyed look at what’s genuinely new and what’s marketing.
The ARC-AGI Score Everyone Is Talking About
One of the loudest claims is that GPT-6 scored close to 100% on the ARC-AGI 3 benchmark. ARC-AGI tests abstract pattern recognition — think simple puzzle games where the rules are never explained to you. You have to infer them through trial and error, then act on them. It’s the kind of thing most humans find trivially easy after a few attempts but that has historically humbled AI models.
Previous frontier models scored around 5 points on ARC-AGI 3. GPT-6 nearly cleared the board. That’s a real leap.
There’s a catch, though. OpenAI used a custom inference harness — essentially a tailored scaffolding layer that handles memory and context differently than the default setup. When the standard harness was used, scores dropped to around 60. The benchmark-busting number came from OpenAI’s own optimized wrapper, not the model running out of the box.
The lesson isn’t that GPT-6 isn’t impressive. It is. The lesson is that how you scaffold an AI matters as much as the model itself — sometimes more.
Computer Use: Keyboard and Mouse Control, No Plugins Required
GPT-6 ships with a significantly improved version of Computer Use, the feature that lets the model directly operate your keyboard and mouse rather than calling a structured API. This is a bigger deal than it sounds.
Previous approaches to AI-driven automation required you to install connectors, write integration code, or use purpose-built tools for each application. With direct input control, GPT-6 can open a spreadsheet, fill in values, switch to a CAD tool, adjust a dimension, and save the file — without a single plugin.
Practical examples that have already surfaced:
- Giving it a reference photo and having it reproduce the image in a paint application, freehand
- Driving Excel workflows that involve conditional formatting and multi-sheet logic
- Operating desktop design tools without any API access
The tradeoff is real: fine-grained precision is still a weak point, and token consumption is higher because the model is processing screen state continuously. For most tasks, a structured integration is still more efficient. But for one-off jobs in software that has no API, Computer Use is genuinely useful right now.
Style Consistency Is Quietly the Best New Feature
Under less fanfare, GPT-6 has gotten noticeably better at maintaining visual and structural consistency across a document or project. If you provide a single slide as a reference — colors, typography, layout density — and ask it to build out a full deck, it holds the theme across every subsequent slide without drifting.
The same applies to written documents: provide a sample section with a specific tone and formatting style, and the model replicates it reliably through a long piece.
For anyone doing client deliverables, internal reports, or templated content at scale, this is the update worth paying attention to.
Coding Benchmarks: Temper Your Expectations
Here’s where the excitement needs a cold splash of water. On standard coding benchmarks, GPT-6’s improvements are modest. It handles agentic tasks — longer, multi-step coding workflows — somewhat better than competing models, and hallucinated result reporting has apparently dropped by a factor of four, which matters a lot when you’re running autonomous agents overnight.
But for the use case most professional developers actually care about — reliably adding features and fixing bugs inside a large, existing codebase — GPT-6 doesn’t show a dramatic edge over lighter, cheaper models. If a flash-tier model handles your real-world repo work just as well and costs a tenth of the price, the economics don’t favor defaulting to GPT-6 for everything.
Use it where it earns its keep. Don’t use it as your default just because it’s new.
A Note on the AGI Declaration
OpenAI has said, in various ways, that GPT-6 represents a meaningful milestone toward AGI. It’s worth remembering that the definition of AGI being applied here is shifting. Scoring well on a benchmark designed to test novel problem-solving is meaningful. It’s not the same as general intelligence.
The ARC-AGI creators themselves are clear that benchmark performance isn’t a sufficient condition for AGI. A model that aces a structured test under optimal scaffolding and then costs the equivalent of $13,000 in compute to clear a simple game isn’t operating like a human brain — which solves the same task on less electricity than it takes to power a dim light bulb for a second.
That efficiency gap is still enormous, and it matters for how you think about deploying this technology at scale.
How to Actually Try It
GPT-6 is currently rolling out to paid subscribers. If you want to experiment without committing to a monthly plan, it’s available through API access on platforms like OpenRouter, where you pay per token.
Start with the tasks where GPT-6’s specific strengths apply: novel reasoning problems, long-form document creation with strict style guidelines, and Computer Use workflows involving software with no programmatic interface. Measure the results against what you were getting before.
Skip the three.js demo. Build something you’d actually use.