Fun Test on the Most Powerful AI Models: Fable 5, Sol 5.6, Kimi K3 and Gemini 3.1 Pro
Fun game test on Fable 5, Sol 5.6, Kimi K3 and Gemini 3.1 Pro, and my experience with these models.

I finally have time to show something fun after doing a few tests on these latest AI models.
For this test I’m gonna show here, I asked each model to create the same browser game from a single prompt: a survival game set inside a damaged Mars research base, where the player has 90 seconds to align solar panels, manage power between life support, heating and communications, and transmit an SOS before a huge dust storm arrives. Everything had to be built in one self-contained HTML file, with no external libraries or assets.
Of course, this is only one very small test. Which model you should use also depends on its other capabilities (research, coding etc.) and most importantly, its consistency and stability.
That is also why I have been looking for a model to replace Claude. I think Claude has too many glitches and lacks consistency, possibly because it is constantly running under heavy demand.
Anyway, based on my own experience:
GPT
GPT is still the most human-like model. It is also the most “street-smart” and probably the best at pretending to be human or replicating human feelings and reactions.
When I need something fast, clean, clear, concise and coherent, ChatGPT is usually the best model to use. It also has much fewer practical usage restrictions, so it is more suitable for heavy-lifting workflows. It will probably remain my first go-to model for almost anything.
Claude
Now, Claude is obviously well known for being more academic and research-focused. Its biggest advantage is the depth and technicality of its research, as well as its focus on trustworthiness. It tends to source, detect and verify information more carefully, and sometimes it is even the model reminding me that I should do the same. It is a very data-led model.
However, it still hallucinates, especially when using anything below Opus (never use Haiku). The information may be more trustworthy than what you get from some other models, but it is still nowhere near 100%.
And that leads to my biggest annoyance. You have to use Claude higher models for an important research or website-building task, but it operates under strict usage limits. At the same time, the usage tracking and display are not always reliable.
Half the time, the AI stops in the middle of a task because the usage limit has been reached, without giving any proper warning in advance. Then, once the limit refreshes, I sometimes have to start again and lose all the credits already spent on the unfinished attempt.
After a few months, I realised I had wasted quite a lot of money on Claude simply because of its limited capacity and the lack of transparency around its usage restrictions.
The same applies to Claude Code. In my opinion, the lower models are not particularly smart, but they can still be quite costly and less consistent than Codex, which often leads to excessive token burning.
Gemini
I have not used Gemini Pro that much. I also do not have many use cases involving image or video generation, which I believe is one of Gemini’s main strengths. It also does not currently offer the same kind of coding-agent experience as some of the other models, so I am still not 100% sure how to judge it.
Kimi
As for Moonshot’s Kimi, I have only been using it for a few days. Based on the use cases they have published, its capabilities seem very strong, and I think a lot depends on how you work with and utilise the model.
However, on the same task I gave all of them, Kimi took around twice as long as Sol 5.6 and around 1.5 times as long as Fable 5, while producing roughly the same or slightly lower quality.
Still, the use cases look very promising. The model is cheaper and seems more consistent than Claude, so I am going to keep spending time with it to see whether it can finally help me get rid of Claude. :)






