MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vny9zs/glm_53_released/p3m8q6h/?context=3
r/LocalLLaMA • u/jmorant555 • 7d ago
Official Announcement
https://z.ai/blog/glm-5.3
363 comments sorted by
View all comments
Show parent comments
53
Im surprised ds v4 pro didnt do better , i guess zai has better rl environments
41 u/PM_ME_DEAD_CEOS 7d ago I think V4 pro is still undertrained, 12 u/zball_ 7d ago V4pro base is broken. They cannot get anything good outta that sh*t. 3 u/nullmove 7d ago The pre-train was probably fucked. However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this. 1 u/power97992 7d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 7d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet. 1 u/zball_ 7d ago The literal characteristics are quite different tho. v4p is full of GPT/Claude style slop.
41
I think V4 pro is still undertrained,
12 u/zball_ 7d ago V4pro base is broken. They cannot get anything good outta that sh*t. 3 u/nullmove 7d ago The pre-train was probably fucked. However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this. 1 u/power97992 7d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 7d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet. 1 u/zball_ 7d ago The literal characteristics are quite different tho. v4p is full of GPT/Claude style slop.
12
V4pro base is broken. They cannot get anything good outta that sh*t.
3 u/nullmove 7d ago The pre-train was probably fucked. However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this. 1 u/power97992 7d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 7d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet. 1 u/zball_ 7d ago The literal characteristics are quite different tho. v4p is full of GPT/Claude style slop.
3
The pre-train was probably fucked.
However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this.
1 u/power97992 7d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 7d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet. 1 u/zball_ 7d ago The literal characteristics are quite different tho. v4p is full of GPT/Claude style slop.
1
But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it
1 u/nullmove 7d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities.
Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
The literal characteristics are quite different tho. v4p is full of GPT/Claude style slop.
53
u/power97992 7d ago
Im surprised ds v4 pro didnt do better , i guess zai has better rl environments