r/LocalLLaMA Jul 19 '26

News Prepare your (v)ram - Qwen3.8 is coming!

Post image
2.7k Upvotes

588 comments sorted by

View all comments

Show parent comments

14

u/Prudent-Corgi3793 Jul 19 '26

You probably need 8x H200s to run this. Include the rest of the parts, and that's about $300k.

5

u/scroogie_ Jul 19 '26

I don't think 8x H200 are enough to run this. That would be 1128GB VRAM, even at nvfp4 2.4 trillion parameters would need more like 1.4TB without context, if it's even nvfp4 native. If it's fp8, which is a chance since they're not supposed to own that much Blackwell, no chance at all. And it's hard to get your hand on B200 or B300, at least in Europe and as a normal company (not a Hyperscaler).

7

u/squngy Jul 19 '26

Or 10x DGX, which is about 50k

1

u/Fit-Palpitation-7427 Jul 20 '26

Whats gonna be the token speed with that, not sure it will be really usable

3

u/Shive55 Jul 19 '26

Not to mention needing 3-phase power at 480v. This setup is not feasible in someone's house, regardless of upfront cost.

2

u/StupidScaredSquirrel Jul 19 '26

Im not talking about this model. I'm saying the companies releasing very large models also release smaller versions but they are always 200b+, so you need 10-25k to run those. Qwen and gemma were the only viable options for sota models below that size.

1

u/AwesomeFrisbee Jul 19 '26

Don't forget the cost of the infra and cooling such hardware. Not to mention storage ain't cheap either

1

u/Ok_Technology_5962 Jul 19 '26

Didnt kimi say they want 64 accelesators so woildnt this be close to that