If you looked the other way, you may have missed the stealth preview model that was free on OpenRouter.AI and OpenCode.AI, which was branded "ox-alpha", that was making a bit of a splash. Everyone had figured out that it was a new Z.ai multi-modal, but no one could figure out how they were giving away 5,835,184,092,873 prompt tokens and 94,402,126,602 completion tokens in a single day!
The answer when it launched as GLM-5.3-Flash is that its an very efficient Flash open-weights model that you can get hosted by a US hardware by San Francisco firm OpenCode.AI for an input cost of $0.07 / 1M and output cost of $0.25 / 1M.
Yet they have hidden one dark truth that I can now reveal. The ultimate proof that this model is different. I have hooked it into my personal fork of the rust based apache2 codex, which I prefer, and I asked it to tell me a joke, and it bombed:
In the era of all the models memorising a snake game and "tell me a joke" it made one up on the spot and bombed 💖💚💛❤️
I constantly test new models with “tell me a joke" when I tree out a new harnesses. You get the same set of canned responses as they have all memorised some “dad jokes”. I have never seen one make up a joke that was not funny. Yet GLM-5.3-Flash trying to improvise a joke. And that is something that I have never seem before. A model that wants to be small, fast, energy efficient, and original. What is not to like?
We are talking about the model that Theo.gg puts up there with Claude Opus and OpenAI Sol. It is good at getting on with huge, long-running agentic tasks as Ox Alpha is INSANE
Last night, I had OpenCode.AI coding agent up and running working across all my current projects and repos going flat out. On one massive repo dust off it is still running after twelve hours, having spent $0.02 max on each bug fix todo, and a grant total for $0.47 total.
Love GenAI or hate it, this new mode is no joke. Boom boom!

Top comments (1)
Your clear breakdown of the stealth preview model in GLM‑5.3‑Flash vs Ox‑Alpha really helped me spot the hidden differences. If you think it could help even more devs, consider syndicating it on ZyVOP (zyvop.com) for broader reach.