GTA benchmark · September 2026
Fable 5.1 blew my mind
Fable 5.1 amazed me with its very first prompt: it generated a huge, lively, intricate city that exceeded my expectations. Another 52 prompts over one weekend turned it into a fun game you can play directly in your browser.
- first output
- 1 h 50 min
- first prompt · before tax
- ≈ USD 50
- follow-up prompts
- 52
- GTA benchmark
- #1
In brief
- The first output took 1 hour and 50 minutes and cost about USD 50 before tax.
- After 52 follow-up prompts, it became a large playable browser game with cars, bicycles, motorcycles, police, and missions.
- Fable 5.1 currently ranks first in my GTA benchmark.
- The game still has bugs, but its scope and level of detail are remarkable.
Two new models
Fable 5.1 against GPT-6 Astra
I regularly test new LLMs on two game benchmarks. Each model receives the same assignment, and I compare the result in terms of functionality, playability, fun, and how thoroughly the game is developed.
Two new frontier models arrived in early September 2026: Claude Fable 5.1 from Anthropic and GPT-6 Astra from OpenAI. Fable 5.1 followed the original Fable 5, which had made a strong impression on me—in my tests, it was the first model able to create a simple playable 3D game from a few prompts. Until then, that had been hard to imagine.
Astra also drew considerable attention for its cybersecurity capabilities in connection with a security incident involving Hugging Face.
I had hoped GPT-6 Astra would match Fable 5.1, but in my game benchmarks it did not even overtake the original Fable 5. That does not make Astra a weak model. It was a major surprise in Blender-based object modeling: it created an intricate, visually impressive 3D bungalow.
I first tested both models on the simpler game in which the player controls a dog herding sheep. Fable 5.1's advantage was not as clear there: I ranked it second behind Claude Opus 5. Opus produced an unusually attractive and polished game with eight levels, four dog breeds, music, inventive interactions, and no serious bugs.
Only the GTA benchmark’s extensive prompt revealed Fable 5.1’s true strength—in my tests, exceptionally demanding tasks clearly suit it.
Then came the GTA game
A huge city that actually feels alive
When I moved on to the GTA benchmark, the result genuinely blew my mind. What Fable 5.1 created is brilliant within the context of this experiment: a huge city with many distinct districts that do not feel like one repeated backdrop.
The sheer size of the map surprised me, as did the number of small details Fable 5.1 added to the game. The city has its own character, and I still run into things I had not seen before.
The city looks fairly lively, and its characters interact with both me and their surroundings. If I run someone over, the people nearby panic. If I start a fight, some run away while others come after me to defend the person I attacked. They are not just decoration—they behave like people.
City details
- cars, bicycles, and motorcycles
- two connected multi-level freeway systems
- bars with music, lights, and dancing crowds
- markets, a port, beaches, residential districts, and a tent camp under a freeway
- police roadblocks, spontaneous fights, and drunk people outside bars
A weekend of fixes
The first output was impressive—and badly broken
Of course, it was not free of problems. Cars had visual artifacts, roads often had no asphalt, vehicle headlights did not work at night, and many other things failed to work as intended. I decided to fix them gradually and ended up writing 52 prompts over one weekend.
The result is an entertaining GTA-style open-world game. Its large city reminds me of the first 3D entries in the series, especially GTA III and Vice City—as long as you can look past the simplified character and building models.
Even the final version is not bug-free. This remains a benchmark output, not a finished commercial game. I identified dozens more bugs and did not tune the missions at all. What matters is that the city holds together: you can drive, fight, complete some missions, or simply explore its districts.
53 screenshots
More views of the game
Show 21 more screenshots Hide additional screenshots
Model comparison
Comparison with other models
These numbers describe only these specific runs; they do not say which model is best in general.
Claude Fable 5.1
- 1 h 50 min
- ≈ USD 50 before tax for the first prompt
- 623K context
- 52 follow-up prompts
GPT-6 Astra
- 35 min
- 3 control-fix prompts
- cost unknown—subscription
- smaller, simpler game
Kimi K3
- about 4 h
- USD 89
- 1 corrective prompt
- behind GPT-6 Astra and GLM 5.3 Flash
The first version of the game took 1 hour and 50 minutes and cost approximately USD 50 before tax. GPT-6 Astra worked for 35 minutes; I do not know the price because the run was part of a subscription. It produced a substantially simpler game with a smaller map, which I ranked behind even the original Fable 5.
Kimi K3 spent about four hours on the assignment and cost USD 89, yet ranked behind even GPT-6 Astra and GLM 5.3 Flash. A model such as GLM 5.3 Flash can run locally on some more expensive computers with 256 GB of RAM, such as a Mac Studio.
My conclusion
The work ethic of Opus 5 combined with Fable's ability
I think Fable 5.1 is a brilliant model, but the assignment matters enormously. In my tests it spends more effort checking its work and can take on a larger scope. That does not win every time—it remained behind Opus 5 in the sheep game—but in the GTA benchmark the combination worked exceptionally well.
Claude Opus 5 felt like a hard worker. For the GTA assignment, it built a dense world full of buildings with better textures, but the buildings overlapped and most of the map could not be explored. The original Fable 5 also created a large world, but most of its area was emptier, with less content between the buildings.
Fable 5.1 therefore feels as though it took Opus 5's work ethic and combined it with Fable's ability. This is the first result that made me feel an AI had produced more than a technical demo of a 3D city—it had made a place I enjoyed exploring.
The game still has many bugs. Even so, it is the largest and liveliest playable 3D game any model has produced in my benchmarks so far.
FAQ
Game FAQ
What did Fable 5.1 create?
A large 3D browser-based GTA-style open-world game for the GTA benchmark.
Can the game be played in a browser?
Yes. It launches directly from the Play the game buttons at the beginning and end of the article.
How long did the first output take?
The first generation took 1 hour and 50 minutes.
How much did the first prompt cost?
The first output cost approximately USD 50 before tax.
How many follow-up prompts were used?
The first output was followed by 52 prompts focused on fixes and tuning within the original assignment.
What can you do in the game?
You can do many of the activities familiar from older GTA games: drive any available vehicle, get into police chases, fight with different weapons, and play some missions.
Can you use cheats in the game?
Yes. Type “cheats” to display the full list of cheats or “arsenal” to unlock every weapon.
Can the game be played on a phone?
Yes. The game works on Android and iOS devices and includes touch controls adapted for mobile screens.
What should I do if the game runs very slowly?
Open Settings and first lower Render Scale and Draw Distance. Set Shadow Quality and Light Range to Low. If the game still does not run smoothly, reduce the other graphics settings as well. The image will be simpler and the viewing distance shorter, but performance should improve substantially.
Does the game save my progress?
Yes. The game automatically saves progress every 30 seconds, when leaving the page, and when starting or completing a mission. It also keeps up to 10 completed-mission saves that can be restored through Load Game. Saves remain in the local storage of that browser and device; they do not transfer between devices and are removed if the site's data is cleared. The exact position of the character is not saved — continuing places the player near the next mission.
Does the game still have bugs?
Yes. It is playable and extensive, but it remains a benchmark output rather than a finished commercial game. I did not tune the missions at all and identified dozens more bugs.