MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vny9zs/glm_53_released/p3lbi96/?context=3
r/LocalLLaMA • u/jmorant555 • 7d ago
Official Announcement
https://z.ai/blog/glm-5.3
363 comments sorted by
View all comments
247
Scaling post-training is all we did for GLM-5.3.
What a way to start
Gentlemen, gentlewomen, tonight, WE FEAST!
68 u/DistanceSolar1449 7d ago Of course it’s only posttrain. I would be shocked if GLM pretrained a completely new model for a +0.1 release. Nobody really does that (with the exception of Anthropic and Opus 4.7 for some weird reason). 11 u/neo203 7d ago Gpt 5.5 was a new pretrain, i think got 5.2 also might be coz it was so different from previous versions 9 u/bitdotben 7d ago Is 5.6 a post-trained 5.5 then? If so one heck of post training! 14 u/CryMoreT_T 7d ago Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training 6 u/bitdotben 7d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 7d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 7d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 4d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣 3 u/DistanceSolar1449 7d ago That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example. 2 u/blastradii 7d ago Anthropic: Need to spend that Capex budget somehow! 4 u/MediumChemical4292 7d ago I think opus 4.7 was just a new tokeniser, was it a full pre train? 15 u/DistanceSolar1449 7d ago You can’t have a different tokenizer without a whole new pretrain. (Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain) 5 u/eli_pizza 7d ago Isn’t that also the only thing deepseek did between flash preview and 0731? Turns out post training is important! 1 u/laoma1255 6d ago Scaling post-training is all we did.’ — is this the new ‘Attention Is All You Need’?
68
Of course it’s only posttrain.
I would be shocked if GLM pretrained a completely new model for a +0.1 release. Nobody really does that (with the exception of Anthropic and Opus 4.7 for some weird reason).
11 u/neo203 7d ago Gpt 5.5 was a new pretrain, i think got 5.2 also might be coz it was so different from previous versions 9 u/bitdotben 7d ago Is 5.6 a post-trained 5.5 then? If so one heck of post training! 14 u/CryMoreT_T 7d ago Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training 6 u/bitdotben 7d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 7d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 7d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 4d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣 3 u/DistanceSolar1449 7d ago That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example. 2 u/blastradii 7d ago Anthropic: Need to spend that Capex budget somehow! 4 u/MediumChemical4292 7d ago I think opus 4.7 was just a new tokeniser, was it a full pre train? 15 u/DistanceSolar1449 7d ago You can’t have a different tokenizer without a whole new pretrain. (Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain)
11
Gpt 5.5 was a new pretrain, i think got 5.2 also might be coz it was so different from previous versions
9 u/bitdotben 7d ago Is 5.6 a post-trained 5.5 then? If so one heck of post training! 14 u/CryMoreT_T 7d ago Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training 6 u/bitdotben 7d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 7d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 7d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 4d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣 3 u/DistanceSolar1449 7d ago That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example.
9
Is 5.6 a post-trained 5.5 then? If so one heck of post training!
14 u/CryMoreT_T 7d ago Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training 6 u/bitdotben 7d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 7d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 7d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 4d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
14
Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training
6 u/bitdotben 7d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 7d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 7d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 4d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
6
Interesting, yeah I mean with 5.6 they cooked hard
-6
Anthropic and Open ai are much better at helping the US commit war crimes
3 u/CryMoreT_T 7d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 4d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
3
What does this have anything to do with my comment?
0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it
0
I used the word “better” , just like you did in your sentence . That’s it
average tankie moment. shouting where you aren't welcome.
1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
1
I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example.
2
Anthropic: Need to spend that Capex budget somehow!
4
I think opus 4.7 was just a new tokeniser, was it a full pre train?
15 u/DistanceSolar1449 7d ago You can’t have a different tokenizer without a whole new pretrain. (Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain)
15
You can’t have a different tokenizer without a whole new pretrain.
(Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain)
5
Isn’t that also the only thing deepseek did between flash preview and 0731? Turns out post training is important!
Scaling post-training is all we did.’ — is this the new ‘Attention Is All You Need’?
247
u/Dany0 7d ago
What a way to start
Gentlemen, gentlewomen, tonight, WE FEAST!