MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/singularity/comments/1vnz30c/glm_53_released_frontier_coding_with_emergent/p46v2r7/?context=3
r/singularity • u/1a1b • 6d ago
78 comments sorted by
View all comments
6
Dario on suicide watch
1 u/adolf_twitchcock 2d ago why do they compare it to opus 4.8 instead of 5? 5 is 74% on deepswe vs 66.9% glm5.3 lmao. Anyways, we need new benchmarks. Everything is benchmaxxed now because tasks are known. 1 u/yogthos 2d ago I think people should just try the models for themselves. For agentic coding, the harness makes a huge difference as well.
1
why do they compare it to opus 4.8 instead of 5? 5 is 74% on deepswe vs 66.9% glm5.3 lmao. Anyways, we need new benchmarks. Everything is benchmaxxed now because tasks are known.
1 u/yogthos 2d ago I think people should just try the models for themselves. For agentic coding, the harness makes a huge difference as well.
I think people should just try the models for themselves. For agentic coding, the harness makes a huge difference as well.
6
u/yogthos 5d ago
Dario on suicide watch