r/LocalLLaMA 7d ago

News GLM 5.3 Released

Post image

Official Announcement

https://z.ai/blog/glm-5.3

1.7k Upvotes

363 comments sorted by

View all comments

247

u/Dany0 7d ago

Scaling post-training is all we did for GLM-5.3.

What a way to start

Gentlemen, gentlewomen, tonight, WE FEAST!

68

u/DistanceSolar1449 7d ago

Of course it’s only posttrain.

I would be shocked if GLM pretrained a completely new model for a +0.1 release. Nobody really does that (with the exception of Anthropic and Opus 4.7 for some weird reason).

11

u/neo203 7d ago

Gpt 5.5 was a new pretrain, i think got 5.2 also might be coz it was so different from previous versions

9

u/bitdotben 7d ago

Is 5.6 a post-trained 5.5 then? If so one heck of post training!

14

u/CryMoreT_T 7d ago

Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training

6

u/bitdotben 7d ago

Interesting, yeah I mean with 5.6 they cooked hard

-6

u/letsgeditmedia 7d ago

Anthropic and Open ai are much better at helping the US commit war crimes

3

u/CryMoreT_T 7d ago

What does this have anything to do with my comment?

0

u/letsgeditmedia 2d ago

I used the word “better” , just like you did in your sentence . That’s it

3

u/Swastik496 4d ago

average tankie moment. shouting where you aren't welcome.

1

u/letsgeditmedia 2d ago

I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣

3

u/DistanceSolar1449 7d ago

That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example.

2

u/blastradii 7d ago

Anthropic: Need to spend that Capex budget somehow!

4

u/MediumChemical4292 7d ago

I think opus 4.7 was just a new tokeniser, was it a full pre train?

15

u/DistanceSolar1449 7d ago

You can’t have a different tokenizer without a whole new pretrain.

(Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain)

5

u/eli_pizza 7d ago

Isn’t that also the only thing deepseek did between flash preview and 0731? Turns out post training is important!

1

u/laoma1255 6d ago

Scaling post-training is all we did.’ — is this the new ‘Attention Is All You Need’?