GPT-6 Astra Beats Portal by Itself

OpenAI's new flagship model played through all of Valve's Portal on its own, and one hobbyist caught the whole 24-hour run.

GPT-6 Astra Beats Portal by Itself

An AI walked through the test chambers

Portal is a first-person puzzle game where you shoot two linked doorways onto walls and floors, then fling yourself between them to solve spatial brain teasers. It trips up plenty of human players. So it is genuinely notable that OpenAI's new GPT-6 Astra model played through the entire game on its own, with no human hands on the controls.

An enthusiast going by CozyBlaze ran the experiment and posted the results. The full playthrough took roughly 24 hours of streams. A trimmed highlights reel runs about two hours, with the model's thinking pauses cut out so it is actually watchable.

How the setup worked

The rig is cleverer than it sounds. Astra controlled Portal through MCP, or Model Context Protocol, a standard way for AI models to call external tools, paired with a modified SourcePauseTool. Here is the trick: the game stays frozen while the model thinks. Astra receives screenshots and data on where the player character is standing, works out a plan, then sends an input sequence. The tool unpauses the game, executes those moves, and pauses again.

That stop-start rhythm is why the run stretched to a full day. Over the course of it, Astra made 3,336 tool calls. The headline API cost came to $571.18 in tokens, though CozyBlaze later clarified that a $200 Codex Pro subscription actually covered it. If you want to poke at the inner workings, the Portal Agent code is on GitHub with instructions to try it yourself.

Why it matters

To appreciate the jump, remember that not long ago AI models were losing at Atari 2600 chess. Steering a general-purpose agent through a 3D puzzle game, reasoning about geometry and physics from screenshots alone, is a different tier of task. It is also a small callback to an old OpenAI goal. Back in 2016, the company listed "solve a wide variety of games using a single agent" among its technical aims.

Worth keeping expectations grounded, though. CozyBlaze is refreshingly pragmatic about it, noting that plenty of problems remain and that this run should not be treated as a proper benchmark. It is a demonstration, not a scoreboard. "Watching a general-purpose agent autonomously navigate and make it all the way through the game feels like a small glimpse of that original vision becoming real," they wrote.

What's next

GPT-6 Astra became OpenAI's flagship model earlier this month. The company describes it as "a new generation of intelligence" and claims it leads on computer use, browsing, software engineering, cybersecurity, science, and professional work. Those are OpenAI's own claims, not independently audited figures, so treat the marketing language with the usual pinch of salt.

Still, the Portal run is a tidy illustration of what "computer use" actually looks like when you point it at something playful. An AI that can read a screen, plan a sequence of actions, and carry them out is doing the same core loop whether the target is a puzzle game or a spreadsheet. The next interesting question is not whether a model can finish Portal, but how much cheaper and faster that same general agent gets at the messier, less scripted tasks people actually want automated.

Subscribe to BuzzBelow

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe