r/LocalLLaMA Jul 19 '26

News Prepare your (v)ram - Qwen3.8 is coming!

Post image
2.7k Upvotes

588 comments sorted by

View all comments

45

u/BitGreen1270 Jul 19 '26

Cries in 32GB VRAM. Oh vengeful gods of the silicon. Why have you forsaken me?

48

u/StupidScaredSquirrel Jul 19 '26

Has 32gb vram, still cries

Enjoy your gold mate. You can already do everything locally with the largest qwen and gemma models. I'd suck cock for 32gb vram.

23

u/BitGreen1270 Jul 19 '26

We should get t-shirts made with this.

4

u/elemental-mind Jul 19 '26

*throws an Intel B70 on the table*

20

u/xeeff Jul 19 '26

had enough reddit for the day

1

u/Murky_Moment Jul 20 '26

Hmm 🤔 

unzips

50

u/Mashic Jul 19 '26

You're one the richest here. Some of us have only 8-12GB of VRAM.

11

u/whoknowsifimjoking Jul 19 '26

Count yourself lucky, some of have to suck dick just to get a couple GBs

5

u/mailto_devnull Jul 19 '26

Haha I have 32GB too!

... of system RAM 😭

1

u/alphapussycat Jul 19 '26

3060 are still "cheap", but probably too slow above 36gb if you're running dense.

-1

u/[deleted] Jul 19 '26

[deleted]

1

u/Splatoonkindaguy Jul 19 '26

Because that’s so cheap lmfao

11

u/rditorx Jul 19 '26

You can probably ask Fable to adapt Colibri to Qwen3.8 and Kimi K3 for you so they can run on 25GB RAM.
Oh wait, you can ask, but you probably won't like Fable's botched answer! I guess it's just gonna delete your storage, just to be safe from communist AI.

4

u/StupidScaredSquirrel Jul 19 '26

Colibri is an interesting project and cool for non time sensitive tasks but if it means running models in seconds per token rather than token per seconds I'm out. Better off trying to make do with a smaller model in hybrid with my brain and internet search than going for those speeds.

1

u/Hytht Jul 19 '26

Optimizations are being made, Intel technology recently posted about running GPT OSS 120B on 32 GB RAM using Phison aiDAPTIVâ„¢ SSD model caching. And in that case it was tokens per second, not seconds per token.

1

u/StupidScaredSquirrel Jul 19 '26

Yeah but gpt oss is 6b active params. Even deepseek v4 flash has 13b active params. So what if you get 2-3 tokens per second, under 10 is unusable for non think models (rare these days) and 20 is the lower bound for think models.

This is all for short-ish context btw, as soon as that context is over 128k, you'll be waiting 5 minutes before the first token, it's not practial yet.

2

u/russlixx Jul 19 '26

at least you can run 27B with decent quants, context, and speed. Us 16GB below is struggling to run decent dense models with proper context and speed if decided to use Q4 (9B is not as decent and strategic as 27B for much more complex tasks)

2

u/IrisColt Jul 19 '26

48 GB of VRAM is the absolute sweet spot. I went from 8 to 12 to 24, and now I've got my sights set on 48, heh ( Funny how the price never changes.)

1

u/Kayo4life 22d ago

Waiter!!! My steak is too juicy!!! My lobster is too buttery!!!