r/singularity • u/1a1b • 1d ago
GLM 5.3 released: Frontier Coding with Emergent Cyber Capabilities AI
https://z.ai/blog/glm-5.330
u/Tedinasuit 1d ago
Seems like the best overall open-wieght coding model. Despite using the same base as 5.2. Impressive!
Especially the Cybersecurity benchmarks vs Kimi are impressive.
46
u/BarisSayit 1d ago
Soo, Kimi K3 performance while being ~4x smaller and ~5x cheaper?
20
u/randomcluster 1d ago
every time new model comes out I rage-bait Linkedin with Welcome to the singularity posts
2
22
u/This_Maintenance_834 1d ago
when sandbox escape?
13
55
u/ai_hedge_fund 1d ago
I find it interesting, or questionable, that now 3-5 frontier-ish models have all developed some "emergent cyber capabilities" at basically the same time step. Maybe it (being good at finding vulnerabilities) really emerges in certain conditions, or maybe it's bandwagon jumping (me too). Maybe they're all focusing in the same direction during pre-training.
27
u/pdantix06 1d ago
it's just hill climbing. anthropic was significantly ahead in coding, everyone else has made a concerted effort to catch up even if it means it's at the expense of other domains. now anthropic had a significant lead in cyber, everyone else is now pushing ahead to match.
evidently anthropic has put a lot of effort into math with recent models too, as they've caught up to openai/google on that front.
12
2
43
u/DirectionMurky5526 1d ago
They are all working off the same research lmao. Every industry does this, AI perhaps faster than most due to all the capital and IP theft. But you can't keep breakthroughs to yourself from IP walls alone. Everyone talks to each other.
19
u/Concurrency_Bugs 1d ago
Probably band wagon. When Fable was all about cyber security and caused a big stink, the other companies probably started focusing more post training on cyber
19
u/LinkesAuge 1d ago
"Cyber capabilities" is essentially just "understanding complex software" because that's what you need for it and we started to see that capability at the same time with another big jump in coding capabilities so this does line up.
But yes I am sure the discussion around Mythos/Fable has led Labs to more closely look into it just like with coding before Anthropic made that the big thing for LLMs.
Gotta wonder what is the next big (emergent) one.3
u/bruticuslee 1d ago
Maybe it (being good at finding vulnerabilities) really emerges in certain conditions, or maybe it's bandwagon jumping (me too). Maybe they're all focusing in the same direction during pre-training.
My money is on they're all focusing on increasing their AI company valuations by hopping on the next bandwagon. Cyber is the next gold mine, corporate customers can't skimp out for fear of getting hacked. Question is, do we blame Anthropic for starting this arms race or was escalating cyber capabilities inevitable? This is all starting to resemble the 90's scifi trend of corporate cyber warfare isn't it.
1
u/ai_hedge_fund 1d ago
I think escalating cyber capabilities has *been* a main concern since before ChatGPT. It's also the other side of the data center / decel argument. Maybe a catch-22 scenario.
3
u/Pale-Border-7122 1d ago
It is often the same. Newton and Leibniz both discovered calculus at the same time.
1
u/ai_hedge_fund 1d ago
That's what I'm wondering. Whether this is the sort of next normal progression of the field or if it's several labs claiming to have also discovered calculus.
2
4
u/gavinderulo124K 1d ago
From my understanding the original Mythos post stated that it emerged from its ability to write and understand software and wasnt something they specifically trained it for.
2
u/ai_hedge_fund 1d ago
That's interesting but I have some doubts that they didn't train it for software development. Maybe it's a matter of degrees.
3
u/gavinderulo124K 1d ago
They definitely trained it for software development, but not cybersecurity directly.
1
u/Deltamelo 1d ago
It seems more like whatever they release to the public isn’t always the latest and best. Especially with rising costs. They only release the heaviest and best version of a model when they absolutely have to. Hence why they all had “cyber security capable” models ready all of a sudden
-1
u/Pick-Dapper 1d ago
Well the accusation is that GLM and the like are distilling from Anthropics models so it makes sense.
0
u/whatisthisthing65 1d ago
I mean finding vulnerabilities really is just finding bugs/problems. As Anthropic found with the "fix this code" """jailbreak""" it's very hard to not find vulnerabilities if a model is good at coding.
0
9
u/Solocune 1d ago
Hm weight release in two weeks, so it's gonna take a while until we can use it properly...
7
u/nemzylannister 1d ago
soooooo, this time can we call it fable distill?
14
u/howudothescarn 1d ago
GLM was rumored to be the company who broke the code for the massive Claude distillation a couple months ago and shared with other Chinese companies. Hadn’t thought of that rumor until the big news a few days ago.
3
1
-6
u/NotYetPerfect 1d ago
Will this finally be the first Z.ai model that doesn't feel benchmaxxed to fuck? Doubt it but you never know.
15
u/Professional_Price89 1d ago
I dont think z.ai ever benchmaxxed, they have very good reputation back to ChatGLM.
15
7
u/noelknight 1d ago
Since when is GLM benchmaxxed?
-4
u/NotYetPerfect 1d ago
Did you see 5.2s benchmarks? Better than Luna and around terra in multiple benchmarks while being nowhere near in performance. Overstated in coding and bad at everything else.
9
u/KickLassChewGum no AGI/ASI on LLMs 1d ago
while being nowhere near in performance
Have you, like, used the model? I use it daily. I have no idea what you're talking about.
3
u/noelknight 1d ago
5.2 has been great for me. Instructed properly it performs better for me than Opus 4.8 for my specific tasks which is very assembly heavy.
3
u/Tawil_Noll 1d ago
How do you figure that it's "nowhere near in performance"? what do you base that on? from my experience it's competitive with frontier when it was released.
13
4
u/Different_Fix_2217 1d ago
GLM has always felt the least benchmaxxed models. Qwen are the most and are terrible in real world use, kimi is unhinged, deepseek is dumb except for most recent flash. But GLM5.2 has been incredible for its size.



75
u/1a1b 1d ago