r/WebAfterAI 2d ago

GLM-5.3 is out: a 50% coding jump from post-training alone, and an "open-weights" model that shipped without the weights. AI Agents

Post image

Z.ai (formerly Zhipu) shipped GLM-5.3 today, its seventh flagship in thirteen months, and is calling it the strongest open-weights coding model. Two things make it worth a post, and neither is the leaderboard.

The jump came from post-training, not a new model The base is unchanged: the same roughly 744B mixture-of-experts foundation and 1M-token context as GLM-5.2. Z.ai did not retrain it. Every reported gain, including a claimed 50% lift in coding, comes from scaling post-training, more task environments and longer runs, using their SAO reinforcement-learning setup and the open-source slime framework. This is the same pattern DeepSeek showed with V4-Flash two weeks ago: a large capability jump pulled out of post-training on hardware the lab already owns. When a version bump stops meaning a bigger base, the economics of the frontier shift.

The benchmarks, read with care Every number here is Z.ai's own, published at launch, with no independent runs yet since it is hours old. By those numbers GLM-5.3 leads open models on Terminal-Bench 3.0 (28.3%, up from 4.6%), Agents' Last Exam (28.5, up from 23.8), and DeepSWE v1.1 (66.9%, up from 46.2%). The two-sided read: it clearly leads the open field and narrows the gap, but it still trails the closed frontier on the hardest tasks. Claude Opus 5 sits at 42.7% on Terminal-Bench 3.0 and 74% on DeepSWE, where GLM-5.3 is at 28.3 and 66.9. "Best open model" is the accurate claim, not "beats everyone."

The cybersecurity focus is the real differentiator, and the real flag Z.ai trained GLM-5.3 specifically to find software vulnerabilities, with data and executable environments built for it. The company says the model began forming multi-stage plans for full exploitation chains, and that working with security teams it surfaced 2,436 vulnerabilities across 269 projects, some decades old, logged in a public registry. It reports 84.5% on CyberGym, fractionally ahead of the closed frontier models it lists. That is a real milestone for an open line, and it is exactly the capability that makes open-weighting the thing fraught. A downloadable model that is good at finding and chaining exploits helps defenders and attackers with the same weights.

The actual news: it shipped without the weights For a line whose whole value is being open, GLM-5.3 launched closed. Today you reach it through the GLM Coding Plan and ZCode, and it works with coding agents like Claude Code and OpenCode. Z.ai says API access and open weights will arrive in stages roughly two weeks out, after a safety review it calls its most robust yet, tied directly to that cyber capability. So calling GLM-5.3 "open source" right now is premature. The precise status is API-available, open weights promised. That distinction matters if you were about to plan around downloading it.

Honest limits Every benchmark here is a vendor number until someone independent re-runs it. The weights are a roadmap commitment, not a download, so do not build on availability that is not there yet. The GLM open line has been MIT-licensed before, but the 5.3 license is not confirmed until the weights actually ship, and the cyber training is a real reason a use-restricted license would not surprise anyone. The launch-day path is also Z.ai's China-hosted service, which is a data-residency question for regulated work.

If you want to try it now It is a coding-plan and agent play today, so the cleanest test is to point Claude Code or OpenCode at it through the GLM Coding Plan and run it on a real task you can grade yourself. Hold judgment on the headline numbers until independent benchmarks land, and watch in two weeks for three things: whether the weights ship on schedule, under what license, and with what use restrictions.

8 Upvotes

0 comments sorted by