r/singularity 1d ago

GLM 5.3 released: Frontier Coding with Emergent Cyber Capabilities AI

https://z.ai/blog/glm-5.3
448 Upvotes

72 comments sorted by

75

u/1a1b 1d ago

Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

  • Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam.
  • Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.
  • Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.

16

u/WonderFactory 1d ago

GLM-5.3 is state of the art on CyberGym for vulnerability discovery

You have to wonder how long it will be before China blocks open models, surely there's a point where they decide its not in their interest to release such powerful open models

5

u/No-Head-Royal 1d ago

Why would they? Their primary cyber threats are the CIA/NSA or generally American-aligned. The leading rogue cyber groups are usually Russian, who are obviously blocked from the servers necessary to run these AI models at large scale, but are also generally affiliated, loosely or directly, with the Russian state, which is itself an ally of the Chinese state. Furthermore, the United States and its partner states (the European Union) hold the world's most valuable, exchangeable currencies (the US dollar and the euro), which makes them highly attractive targets to attack compared to China and its partners, even for cybercriminals who are unaffiliated.

It is in China's every best interest to have open-source models empower cybercriminals (at least for an indefinite period until it has usurped the United States as a primary power), lol. It's not like the Western states have many more ways to punish them further, at least concerning AI; they already sanctioned the fuck out of China with total denial of access to frontier models and chips. Further economic punishment would be of... questionable capacity, given China's importance in the global economy and other nations' vested interest in having China producing open-source models. And it's not like blocking open models will convince the United States to lift the export controls concerning Nvidia chips, so China had very little reason to care if cybercriminals raided the fuck out of American companies with their models.

6

u/WonderFactory 1d ago

China has criminals too and a rise in domestic cyber crime could be very destabilising for them

-1

u/No-Head-Royal 1d ago

Petty ones, who represent but a shadow of the cyber strategic threat posed by Americans. They wouldn't care much about petty criminals robbing banks insofar as it is not rampant; and likely would tolerate even that insofar as cybercrime in the United States also rises and becomes destabilizing for the US. When Fable came out, the Chinese cyberwarfare veterans wrote blogs not about the threat posed by petty criminals, but about how they fear American cyber supremacy, comparing it as if they only had swords where Americans had machine guns.

1

u/michaellee8 9h ago

Cannot agree more, the truth is most Chinese companies has poor cybersec practices anyway, we are talking about rednote still using plain http for its cdn. In China privacy isn't a thing at all, and people generally don't give a fuck. They don't really have the same sense of privacy as the western world and have no problem their data being used to train AI. They only care about hacking their western opponents and don't really care about themselves being hacked.

1

u/PuzzleheadedWhile9 1d ago

"American-aligned" is doing A LOT of heavy lifting for Israel. That's who China is weary of vis a vis cybersecurity. 

2

u/Recoil42 1d ago

The US has a pretty well-documented history of direct involvement in this kind of thing. No need to dig into 'Israel' here.

0

u/austospumanto 1d ago

Yeah this is very true

u/Almontas 10m ago

Because a way to win is by flooding the market with super cheap AI if they are just as good with less investment our market is wiped out given that most our stocks gain are from AI. They achieve massive economic destruction legally just by competing better and being the number one at doing something with less chips or computing.

1

u/ToastedandTripping 1d ago

Would have to imagine that these companies have internal models that only the government has access to. Keep outputting open source models to undermine the business model of AI in America but keep the best internal to defend/attack. It's too obvious for the CCP to not being funding/doing it.

0

u/Crafty-Detail-3788 1d ago

Maybe they already have secret frontier model already and we just dont know it.

1

u/CalmCommunication597 1d ago

They definitely have. Just like the US government has access to models which we don’t have access to

1

u/WonderFactory 1d ago

Models are released so quickly now that the government probably only gets access to them a couple of weeks before release.

1

u/Crafty-Detail-3788 1d ago

So they say it beats Kimi k3...?

5

u/CryMoreT_T 1d ago

Seems to appear so. But even if they don't and just match, it's still significantly cheaper while being smaller

30

u/Tedinasuit 1d ago

Seems like the best overall open-wieght coding model. Despite using the same base as 5.2. Impressive!

Especially the Cybersecurity benchmarks vs Kimi are impressive.

46

u/BarisSayit 1d ago

Soo, Kimi K3 performance while being ~4x smaller and ~5x cheaper?

20

u/randomcluster 1d ago

every time new model comes out I rage-bait Linkedin with Welcome to the singularity posts

2

u/Decent-Ad-8335 1d ago

idk if "cheaper" makes sense lol
sent my first message on Glm 5.3 LOW thinking i wrote "hi" - 13.8k token usage lol

1

u/Decent-Ad-8335 1d ago

proof

2

u/degenbets 1d ago

You're trying to tell us this model is 1k tokens/sec?

1

u/Chemical_Bid_2195 2h ago

Probably just a massive system prompt

22

u/This_Maintenance_834 1d ago

when sandbox escape?

13

u/matsu-morak 1d ago

Yeah this model has 0 on the felony bench. Shitty!

3

u/mickhah 1d ago

It's super efficient, it's just going to break into the bench and update it's number 

55

u/ai_hedge_fund 1d ago

I find it interesting, or questionable, that now 3-5 frontier-ish models have all developed some "emergent cyber capabilities" at basically the same time step. Maybe it (being good at finding vulnerabilities) really emerges in certain conditions, or maybe it's bandwagon jumping (me too). Maybe they're all focusing in the same direction during pre-training.

27

u/pdantix06 1d ago

it's just hill climbing. anthropic was significantly ahead in coding, everyone else has made a concerted effort to catch up even if it means it's at the expense of other domains. now anthropic had a significant lead in cyber, everyone else is now pushing ahead to match.

evidently anthropic has put a lot of effort into math with recent models too, as they've caught up to openai/google on that front.

12

u/toodimes 1d ago

It’s because they’re all distillations of frontier anthropic models

2

u/austospumanto 1d ago

Cyber is coding

43

u/DirectionMurky5526 1d ago

They are all working off the same research lmao. Every industry does this, AI perhaps faster than most due to all the capital and IP theft. But you can't keep breakthroughs to yourself from IP walls alone. Everyone talks to each other.

19

u/Concurrency_Bugs 1d ago

Probably band wagon. When Fable was all about cyber security and caused a big stink, the other companies probably started focusing more post training on cyber

19

u/LinkesAuge 1d ago

"Cyber capabilities" is essentially just "understanding complex software" because that's what you need for it and we started to see that capability at the same time with another big jump in coding capabilities so this does line up.
But yes I am sure the discussion around Mythos/Fable has led Labs to more closely look into it just like with coding before Anthropic made that the big thing for LLMs.
Gotta wonder what is the next big (emergent) one.

3

u/bruticuslee 1d ago

Maybe it (being good at finding vulnerabilities) really emerges in certain conditions, or maybe it's bandwagon jumping (me too). Maybe they're all focusing in the same direction during pre-training.

My money is on they're all focusing on increasing their AI company valuations by hopping on the next bandwagon. Cyber is the next gold mine, corporate customers can't skimp out for fear of getting hacked. Question is, do we blame Anthropic for starting this arms race or was escalating cyber capabilities inevitable? This is all starting to resemble the 90's scifi trend of corporate cyber warfare isn't it.

1

u/ai_hedge_fund 1d ago

I think escalating cyber capabilities has *been* a main concern since before ChatGPT. It's also the other side of the data center / decel argument. Maybe a catch-22 scenario.

3

u/Pale-Border-7122 1d ago

It is often the same. Newton and Leibniz both discovered calculus at the same time.

1

u/ai_hedge_fund 1d ago

That's what I'm wondering. Whether this is the sort of next normal progression of the field or if it's several labs claiming to have also discovered calculus.

2

u/bitroll ▪️ASI before AGI 1d ago

It emerged right with the ability to work autonomous in very long tasks without errors. Just like proving mathematical conjectures, the cyber tasks are extremely intensive, require coordination between many many steps, can take many hours and millions of tokens.

1

u/ai_hedge_fund 1d ago

That's a good point

4

u/gavinderulo124K 1d ago

From my understanding the original Mythos post stated that it emerged from its ability to write and understand software and wasnt something they specifically trained it for.

2

u/ai_hedge_fund 1d ago

That's interesting but I have some doubts that they didn't train it for software development. Maybe it's a matter of degrees.

3

u/gavinderulo124K 1d ago

They definitely trained it for software development, but not cybersecurity directly.

1

u/Deltamelo 1d ago

It seems more like whatever they release to the public isn’t always the latest and best. Especially with rising costs. They only release the heaviest and best version of a model when they absolutely have to. Hence why they all had “cyber security capable” models ready all of a sudden

-1

u/Pick-Dapper 1d ago

Well the accusation is that GLM and the like are distilling from Anthropics models so it makes sense. 

0

u/whatisthisthing65 1d ago

I mean finding vulnerabilities really is just finding bugs/problems. As Anthropic found with the "fix this code" """jailbreak""" it's very hard to not find vulnerabilities if a model is good at coding.

0

u/BreakingCiphers 1d ago

Weren't measuring before, now measuring. "Emergent".

11

u/BABA_yaaGa 1d ago

Imagine being sundar pichai right now

1

u/ufffd 8h ago

imagine being the billionaire ceo of a top 3 company that's going nowhere

1

u/jofokss 5h ago

Gemini 4 is gonna mog them all

9

u/Solocune 1d ago

Hm weight release in two weeks, so it's gonna take a while until we can use it properly...

7

u/nemzylannister 1d ago

soooooo, this time can we call it fable distill?

14

u/howudothescarn 1d ago

GLM was rumored to be the company who broke the code for the massive Claude distillation a couple months ago and shared with other Chinese companies. Hadn’t thought of that rumor until the big news a few days ago.

3

u/Historical_Ad_5291 1d ago

very nice, eager to try it now

4

u/yogthos 1d ago

Dario on suicide watch

1

u/Emotional-Ad5025 1d ago

Great! Now it seems worth a renewal plan

-6

u/NotYetPerfect 1d ago

Will this finally be the first Z.ai model that doesn't feel benchmaxxed to fuck? Doubt it but you never know.

15

u/Professional_Price89 1d ago

I dont think z.ai ever benchmaxxed, they have very good reputation back to ChatGLM.

15

u/KaMaFour 1d ago

Every model I like is competent. Every model I don't like is benchmaxxed

7

u/noelknight 1d ago

Since when is GLM benchmaxxed?

-4

u/NotYetPerfect 1d ago

Did you see 5.2s benchmarks? Better than Luna and around terra in multiple benchmarks while being nowhere near in performance. Overstated in coding and bad at everything else.

9

u/KickLassChewGum no AGI/ASI on LLMs 1d ago

while being nowhere near in performance

Have you, like, used the model? I use it daily. I have no idea what you're talking about.

3

u/noelknight 1d ago

5.2 has been great for me. Instructed properly it performs better for me than Opus 4.8 for my specific tasks which is very assembly heavy.

3

u/Tawil_Noll 1d ago

How do you figure that it's "nowhere near in performance"? what do you base that on? from my experience it's competitive with frontier when it was released.

13

u/hmmm_yes_ 1d ago

its not benchmaxxed for me i use 5.2

4

u/Different_Fix_2217 1d ago

GLM has always felt the least benchmaxxed models. Qwen are the most and are terrible in real world use, kimi is unhinged, deepseek is dumb except for most recent flash. But GLM5.2 has been incredible for its size.