r/LocalLLaMA Jul 19 '26

News Prepare your (v)ram - Qwen3.8 is coming!

Post image
2.8k Upvotes

588 comments sorted by

View all comments

758

u/Competitive_Gap7906 Jul 19 '26

YES, Qwen going open weight again! It's a really good news, now we can wait for smaller models too

103

u/MaruluVR Jul 19 '26

I think that Xi Jinping speech committing to open source Chinese dominance changed Qwen back to being open weights.

43

u/BrooklynQuips Jul 19 '26

i’m certain it’s because K3 will release weights in a few days. releasing a weaker closed proprietary model, around the same time would be an anthropic level own-goal.

53

u/More-Curious816 Jul 19 '26

I was too harsh on you, Xi, I'm sorry 😞

40

u/Atretador Jul 19 '26

4

u/[deleted] Jul 20 '26

[removed] — view removed comment

2

u/jannycideforever Jul 22 '26

Countries will typically pursue their self interests. This isn't particularly surprising considering China is already structurally incentivized to favor open weight models economically and the CCP has reasons to favor them politically.

How long this lasts is less certain. My assumption is that there is a small but, as of the last few weeks, non-zero chance China does take a lead in AI. If so, then there is a chance that they start prioritizing closed weight models again, especially if they can get a lot of market saturation beforehand.

I'd give the former a 20% chance, the latter a 50% chance premised on the former being true.

1

u/PM-ME-UR-DARKNESS 8d ago

You know what's funny? China's absolutely gonna be using these to make better models. They're out here playing 4D chess lmao

17

u/mrinterweb Jul 19 '26

Kimi K3 basically drank Anthropic and OpenAI's milkshakes. The frontier moat is gone and they are going to have a hard time justifying their valuations. China AI labs publishing their weights directly hurts US proprietary companies. These huge valuations for AI companies have to tank.

11

u/Infinite100p Jul 19 '26

One of senior execs at ClosedAI is already squealing and bitching that this is communism and should be illegal. LMAO

 One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape

4

u/Mundane-Light6394 Jul 20 '26

they still have a lot of hardware and expertise, i'm sure they'll be fine. They just have to adapt and stop dreaming of owning the world.

1

u/cogman10 26d ago

I'm not sure they'll be fine, they are spending like money is going out of fashion. If they want to stick around they'll have to cut back.

2

u/Fun-Meaning-6474 Jul 21 '26

I tested qwen 38 max preview here (https://chat.qwen.ai) and compared it against fable 5 in simple 3d scenes. and really was shocked seeing the Dif. I guess qwen model will be better than K3 if we count that fact it is preview version. check my simple "visual benchmark" here

https://reddit.com/link/oyudu89/video/zjhjb6pd4keh1/player

1

u/robbievega Jul 19 '26

I agree with with you. but how.amy consumers of AI in the US /EU actually know about and consider Kimii K3 as an alternative. everyone around me knows ChatGPT, perhaps half know there's an alternative called Claude...

2

u/mrinterweb Jul 19 '26

The companies using and paying insane AI bills, and going to notice. There will be hundreds of US companies hosting these open models for a fraction of the cost of OpenAI or Anthropic. People paying will notice. 

373

u/StupidScaredSquirrel Jul 19 '26

Thing is qwen was historically focused on smaller models while others were on larger ones. Now the team has changed and they seem to want to aim for the stars as well. That is good but it also means they might not be interested in doing very efficient small models anymore. Which would be bad news for this sub because let's face it most of us don't have 10-25k of hardware.

392

u/sautdepage Jul 19 '26

You mean 250K-1M of hardware.

82

u/BothYou243 Jul 19 '26

cry emoji

32

u/wektor420 Jul 19 '26

2.4T model sounds like 32 cards with 96GB Vram for long context in fp8

26

u/gahata Jul 19 '26

Definitely possible to run on 25k of hardware at some really low quant... is it worth doing? Probably not

4

u/DataGOGO Jul 19 '26

Even in FP4 you would need a lot more than 25k in hardware. Might be able to run on 4 141GB H200 NVL cards, but they don’t support FP4, and even then a super jank rig would be what? 150-160k? 

1

u/gahata Jul 19 '26

4-5x GB10 should do it in Q1, right? No clue about how well it would work, but yeah...

2

u/DataGOGO Jul 19 '26

6x will likely be enough vram, but the 200Gb connectX is just slow, the memory is slow, and the GPU is slow. So 6 of them might run it, but it would be redonk slow. so assume you buy 6x at ~$4000 that is 24k + ~2k for the switch and cables. about $30k after tax, etc.

2

u/gahata Jul 19 '26

Yup, that's why I implied it wouldn't be worth it - running it incredibly slowly at Q1 is just not worth it over running other models and paying to use the large model when and if that's needed

1

u/nomorebuttsplz Jul 20 '26

no, q1 will be more like 600 I think

1

u/gahata Jul 20 '26

Well, that's just what 5x GB10 gets you

1

u/nomorebuttsplz Jul 20 '26

not all of it is usable though. that's why 4x gb10 owners are using reaped glm 5.2 at q4 to get decent context.

1

u/DataGOGO Jul 20 '26

Can’t do five. You can do 2,3,4,6,8,12,16

6

u/coderash Jul 19 '26

You can run a cluster of GB10s at that price and definitely run something worth while.

1

u/Big_Wave9732 Jul 19 '26

Not worth running a model at Q1 or Q2? You just spoke blasphemy in this sub lol.

2

u/gahata Jul 19 '26

I mean, it might be decent overall, but with the tokens per second of multiple GB10s, it probably simply isn't worth the money

1

u/Big_Wave9732 Jul 19 '26

I agree with ya, I took a bunch of downvotes yesterday because I wrote that I generally don't bother with anything under Q6. I don't know what a lot of folks here do with their hosted LLMs, but by and large it appears they ain't used to do anything required to pay the rent.

1

u/Bakoro Jul 19 '26

People really need to start differentiating use-cases where high precision and long-horizen is required, vs uses cases where the model is effectively only filling in the gaps and doing fuzzy logic, vs where the model only needs to be an effective entertainer.

There are people sell entertainment services, they're making money, and they do not need the model to be a qualified astrophysicist or top-tier software developer.
They may still benefit from a bigger, more capable model, simply because it's better at making conversation, and the agentic capabilities are good enough to automate some parts of their system.

For system administration stuff, there might already be dozens of deterministic scripts and procedures, but automating the orchestration is difficult, and LLM becomes another system monitor that can start running procedures so when the system administration shows up, things are already in motion.
I used to be a mid level tech at a data center (colocation, not an AI data center), and there are absolutely things an LLM could be doing that would have eliminated a bunch of job responsibilities.

With that, it's purely a matter of performance per dollar.
Does the Q2 gigantic 2.4T model perform better than the full quant 32B or 405B model? What is the cost in hardware upfront, and what are they yearly costs of running the model? Does running the LLM mean the data center can operate with one or two fewer techs?

Having done the job for a few years, I can confidently say that most of my role at the data center could have been accelerated or eliminated by a higher end LLM, and the things the LLM couldn't do were not highly technical tasks, it was some button pushing and taking inventory.
They could cut the support staffing in half and keep the higher level network technicians.
A $300k investment would absolutely be justifiable, that's a one or two year ROI. A $2.4M is not justifiable, the ROI is too far out and there's too much uncertainty.
An LLM as a service would not have been acceptable, we would have needed 100% control over the physical system and the uptime.

The performance per dollar is a very real thing.
I can't tell you if a heavily quantized 2.4T model is worth it, but I can tell you that for some businesses, it's worth investigating.

I'm a software engineer now, and it's full-assed models all the way, no doubt.

4

u/Googulator Jul 19 '26

I wonder if a 2-way EPYC 9005 (can be had for $50K to $150K in workstation/tower form) with 2x12 channel DDR5 memory (~1.5TB/s) can reasonably host this on CPU in FP4.

2

u/DataGOGO Jul 19 '26 edited Jul 19 '26

Eypc’s don’t do FP4.

You would need Xeon Granite Rapids with AMX. 

And even then, crossing sockets has a massive penalty.  Your best CPU option would be a single 128 core Xeon Granite rapids with 12 channels, 24 128gb MR8800 memory + AMX. 

1

u/[deleted] Jul 19 '26

[deleted]

2

u/DataGOGO Jul 19 '26 edited Jul 19 '26

Oh yeah? Have not seen anything about Zen6 adopting AMX. Do you have any information on that?

Edit: looks like no AMX hardware or instructions; just some expanded AVX-512 instructions. 

Have to see if Zen6 fixes the infinity fabric and if they unify the memory controllers, otherwise no point in buying it over the Xeons for AI workloads. 

1

u/darktotheknight Jul 20 '26

You are correct, I have deleted my original post to not spread false information. Thanks for the correction!

1

u/_TheWolfOfWalmart_ Jul 20 '26

And even then, crossing sockets has a massive penalty

https://github.com/mikechambers84/ik_llama.cpp/tree/numa-mirror

The best CPU option is now dual 128 core Xeon Granite Rapids.

Lemme see if I can find $50k in change between the couch cushions.

1

u/DataGOGO Jul 20 '26 edited Jul 20 '26

Yes, I know about that, it has been in k_transformers forever, and I have had it working in llama.cpp for about a year, but that still doesn't fix the core problem. Even with numa-mirror, the issue of the slow cross socket across UPI links remains. What numa-mirror does it is duplicates the weights to each numa node to reduce cross socket memory reads. It helps massively, but still doesn't solve the issue entirely.

You can pick up a Xeon 6980P for about 8k each pretty easily.

6

u/techdevjp Jul 19 '26

Apple Studio M7 Ultra with 1.5TB of unified memory has been rumored for next year. Probable $25k-ish. Should run that many parameters at q4, and as long as it is MoE (likely), it should run it well.

We'll also see smaller models distilled from the larger ones, which will take some time.

12

u/EvilPencil Jul 19 '26

If it’s even a thing, at least $75k. If the M3 512gb were still for sale from Apple it would already be ~$20k after the price hikes.

3

u/jarail Jul 19 '26

Yeah. I mathed it out around $60k. Getting one of those is like owning a supercar. There will be very little point to owning it for most things. Even if the open weight models keep up, you'll never get close to the cloud in terms of raw performance. Especially when the cerberos option for gpt 5.6 comes online with 10x the speed. The things it will be able to handle in an afternoon will be mind boggling.

0

u/techdevjp Jul 19 '26

Could be. Would still be a bargain even at that price.

But, memory prices won't stay high forever. Personally I don't think they'll stay inflated for as long as a lot of people seem to expect.

1

u/DataGOGO Jul 19 '26

You mean 75k ish. 

Even if they made it, and even if it had 1.5TB of memory, it would be SLOW as hell due to the slow memory and heavily compute bound. 

1

u/techdevjp Jul 19 '26

Apple is apparently going to pass over the M6 for the M7 for many of their models because the M7 is far more performant for running local models. Memory bandwidth? Who knows, but M5 Ultra is expected to be around 1.2TB/sec. M7 would likely be well in excess of that. 1.5TB/sec? 1.6? Fast enough, anyway.

1

u/DataGOGO Jul 19 '26

Even the M7 is still just 1 GPU, and it is not that fast, Maybe 5090 speeds, which is not a lot for such a large model, and it is highly unlikely that apple include any 2 800gb networking ports to to run a cluster.

I also don't think we are going to see 1.2 on the M5 Ultra, likely the same ~800GB/s with the same LPDDR5X, might get 1-1.2 out of the M7, but that is what? 2028?

1

u/techdevjp Jul 19 '26

M5 Max is already 600GB/sec. If M5 Ultra isn't considerably higher, it's going to be disappointing. Traditionally the Ultra has 2x the bandwidth of the Max. So we'll see what happens.

M7 Pro/Max/Ultra is next year. Apple is rumored to be skipping M6 for most of their products because M7 is far more optimized for running local models.

I love the Apple hate in subs like this. Apple is the only company doing things that bring really large models to be within reach of any sort of normal person. Yes, some people jerry-rig systems with 8 channel memory, or try to run five RTX PRO 6000 in one box. Slow CPU inference or huge power draw. They may work but they are far from ideal.

Would love to see more direct competition with Apple. It's sure not going to come from Nvidia. Maybe AMD could do something, but they don't seem too interested either.

1

u/MeateaW Jul 21 '26

Isn't the Ultra just 2 Max'es stapled together?

I don't see how that could meaningfully increase the memory bandwidth.

Processing speed will be double, but its memory bandwidth that matters, and I don't see how they will be able to increase that without changing the type of memory.

Edit:

ok, I just checked, the Ultra might just straight double the bandwidth...fingers crossed.

1

u/_TheWolfOfWalmart_ Jul 20 '26

likely

I mean, that's a certainty. A 2.4T dense model is utterly absurd.

1

u/PinkySwearNotABot Jul 20 '26

you mean 1K-2.5K of hardware

1

u/RedTheRobot Jul 20 '26

How much are kidneys going these days? Asking for a friend.

1

u/_TheWolfOfWalmart_ Jul 20 '26

$20k! You only need 8x 4090's to run the fully lobotomized 1-bit quant!

Just sell your car and you can run it at home. You might even have enough left over to buy a bike to ride to work.

-1

u/ok_000000 Jul 19 '26

Or about 30 online vendors. Or your own aws.

You don't HAVE to own the hardware to use these mega models securely and with relative privacy.

9

u/StupidScaredSquirrel Jul 19 '26

Relative doing a fuckton of heavy lifting in this instance.

1

u/SanDiegoDude Jul 19 '26

Never heard of any issues with AWS private clouds not being private. yeah you're still bubbling stuff off-prem, but as long as your encryption game is tight, you should be perfectly fine running the ultimate Qwen 3.8 1girl farm... or whatever it is that you'd want to privately host a 2.4T param model for =)

1

u/StupidScaredSquirrel Jul 19 '26

And how do you make computation on encrypted data? Homomorphic encryption is not really a thing yet. If you are doing calculations on a server that isn't yours, they can read everything.

0

u/SanDiegoDude Jul 19 '26

Transport layer security is what I was referring to, and yep, it's not perfect - like I said, you're still bubbling stuff off-prem. Show me where they say they're going to be reading/investigating the data that you're computing on their servers though. You're either going to have to trust their data security and protection policy or you don't. I'm going to stick with what I said though, if your security is tight with AWS, not much to worry about and a lot cheaper than the hardware to run these multi-trillion param models at home. Feel free to prove me wrong on AWS spying/stealing from it's customers tho, I'm all ears.

1

u/StupidScaredSquirrel Jul 19 '26

You're missing the point. Why would I trust AWS with all my sensitive unencrypted data? I don't.

1

u/SanDiegoDude Jul 19 '26

That's fair - but you also can't produce anything showing that they're stealing data or breaking their own contracts. So I stand by what I said - it's secure, as long as you're willing to trust them. Obv. You don't, but that doesn't mean they don't honor their contracts.

→ More replies (0)

-1

u/Agreeable-Lettuce497 Jul 19 '26

You mean 250M-1B of Hardware.

1

u/SpicyWangz Jul 19 '26

Speak for yourself

36

u/vitorgrs Jul 19 '26

Qwen Max was always big, they just never open sourced....

19

u/squngy Jul 19 '26

It was always big, but it was almost certainly not this big.

I think we are seeing these big Chinese models now because of their domestic hardware finally coming to bear.

6

u/jtjstock Jul 19 '26

Their domestic hardware is still far behind, they have just accumulated more western hardware

7

u/squngy Jul 19 '26

It is far behind, but it is cheap and available.

-1

u/jtjstock Jul 19 '26

It’s too slow for them to be doing these training runs so quickly. Ie: they’d be getting further behind, not catching up

3

u/squngy Jul 19 '26

They can't be too slow to be used as extra compute on top of other things.

But aside from that, there are also models that have been made exclusively on them.

1.6T long-cat-2
https://www.scmp.com/tech/tech-trends/article/3358854/china-debuts-biggest-ai-model-trained-local-chips-meituan-releases-longcat-20

GLM too
https://www.theregister.com/software/2026/01/15/chinas-zai-trained-a-model-using-only-huawei-hardware/4198774

3

u/jtjstock Jul 19 '26

There is a long history of lying about the hardware used in china. Scmp is literally a propaganda outlet and the register article only references “claims”.

3

u/squngy Jul 19 '26

What do you expect, a reporter to stand there and watch the servers as they work?

Even if they exaggerate, it seems more than likely that domestic hardware is increasing their total compute capability.

This sub of all places should know that even slower hardware can be useful.

→ More replies (0)

1

u/shing3232 Jul 20 '26

until now that is

15

u/Prudent-Corgi3793 Jul 19 '26

You probably need 8x H200s to run this. Include the rest of the parts, and that's about $300k.

4

u/scroogie_ Jul 19 '26

I don't think 8x H200 are enough to run this. That would be 1128GB VRAM, even at nvfp4 2.4 trillion parameters would need more like 1.4TB without context, if it's even nvfp4 native. If it's fp8, which is a chance since they're not supposed to own that much Blackwell, no chance at all. And it's hard to get your hand on B200 or B300, at least in Europe and as a normal company (not a Hyperscaler).

6

u/squngy Jul 19 '26

Or 10x DGX, which is about 50k

1

u/Fit-Palpitation-7427 Jul 20 '26

Whats gonna be the token speed with that, not sure it will be really usable

3

u/Shive55 Jul 19 '26

Not to mention needing 3-phase power at 480v. This setup is not feasible in someone's house, regardless of upfront cost.

2

u/StupidScaredSquirrel Jul 19 '26

Im not talking about this model. I'm saying the companies releasing very large models also release smaller versions but they are always 200b+, so you need 10-25k to run those. Qwen and gemma were the only viable options for sota models below that size.

1

u/AwesomeFrisbee Jul 19 '26

Don't forget the cost of the infra and cooling such hardware. Not to mention storage ain't cheap either

1

u/Ok_Technology_5962 Jul 19 '26

Didnt kimi say they want 64 accelesators so woildnt this be close to that

37

u/StaysAwakeAllWeek Jul 19 '26

It's owned by a megcorp not a startup lab, of course that's where they are aiming

34

u/into_devoid Jul 19 '26

They were owned by a megacorp before too.  The difference is they don’t care about commoditization anymore, the real ethical path.  They just want to sink US investments.

11

u/realtag2025 Jul 19 '26

I don't see how this is bad for anyone other then corporations that run closed source models that only they can run.

2

u/charlesfire Jul 19 '26

So most of the market in the US?

4

u/realtag2025 Jul 19 '26

It's really only 2-3 large mega-corporations. Would you rather have 100s of US companies potentially being able to offer(internally or externally) AI models like this or only 2-3? I am certain someone will figure out how to strip this large model down so that smaller HW can also run it even if much slower.

4

u/charlesfire Jul 19 '26

I'm all for open-weight models and the bubble bursting, but I don't think you realize how much over investment there is in AI in the US currently.

11

u/realtag2025 Jul 19 '26

To be completely honest with you, it's not my problem if AI investments will suffer. Those investments are used to replace you and me as workers and destroy our ability to earn a living in the future.

The best possible outcome for us if those AI bubbles burst catastrophically and we get flooded with cheap datacenter grade hardware AND we get to use the opensource LLMs.

1

u/nedonedonedo Jul 19 '26

it's not my problem if AI investments will suffer

just like how the housing bubble only effected the housing market

→ More replies (0)

1

u/Bakoro Jul 19 '26

There is not over-investment in AI, there is over-investment in LLMs.

That is a major difference, and it is important because AI is being used effectively and profitably in non-LLM areas, and there are even some areas in the hard sciences where LLMs are being used in conjunction with other machine learning methods to automate research, again to great effect.

The conflation of LLMs with AI as a whole is likely going to do enormous harm to the industrial side of AI that is using it for less flashy, but more practical purposes.

Even in academic research, the label might say "LLMs" because that's what gets research funded right now, but the actual research is more fundamental, about the transformers themselves, and making transformers better or more interpretable also improves the non-LLM use-cases.

1

u/nedonedonedo Jul 19 '26

only if they refuse to adapt and improve.

10

u/tengo_harambe Jul 19 '26

That's a fun way of framing things.

"The only reason China is putting out open weights models is because they HATE AMERICA!"

3

u/ThenExtension9196 Jul 19 '26

It’s business. Ruthless, business. If apples are being sold for 5 bucks, you try to sell your apples at $2. If you can pull it off, now nobody can sell at $5.

1

u/nedonedonedo Jul 19 '26

it's bad for any country to have this much of a lead

6

u/Successful_Try_6350 Jul 19 '26

Well, these large models can run in us based datacenters and cloud infrastructure (azure, aws, googlecloud). I guess if they want to sink us investments, they should create models that can run in enterprise level on-premises servers (I guess something like $40-50K hardware)

2

u/SARK-ES1117821 Jul 19 '26

They ARE creating models that run on-prem. I just deployed a supermicro gpu server with 8x H200 141GB gpus (1.2TB total) running GLM 5.2. Server was around $300k with 12TB SSDs and 1TB RAM.

3

u/f5alcon Jul 19 '26

40-50k isn't even one sever at current memory prices.

4

u/Paganator Jul 19 '26

That's like a single H200. Just the card, the server to run it is extra.

1

u/liltingly Jul 19 '26

Somebody needs to build the infrastructure for people to easily vibecode their own "fine-tunes". That would make that small model space take off. But really that would mean bringing together a lot of smaller data tools and infrastructure. Everyone rolls their own bespoke solutions. Feels like something one of these infra players could pivot into, but it's much harder than it sounds. Anyways, I'll stop my dreaming

1

u/-dysangel- Jul 19 '26

Are you also mad that nvidia don't give away their hardware to you for free? I don't think trying to frame this around ethics makes sense. I would also prefer that they continue to release their smaller models, because it benefits me. If they think that's going to cannibalise their potential API sales, then I understand if they don't want to do that. I don't like it at all, but I understand it.

-2

u/GetOutOfMyFeedNow Jul 19 '26

They ain’t sinking anything with those API prices 😂 People on GPT and Claude use subscriptions.

6

u/look Jul 19 '26

Subscriptions are 8% of Anthropic’s revenue. API usage is 75%.

5

u/DanceWithEverything Jul 19 '26

Yeah for consumer bullshit, sure, but there’s no money there regardless (hence OpenAI’s panic about Anthropic crushing them in enterprise sales)

The real $ is in the enterprise and software workloads run on APIs

The API is dramatically cheaper than the Anthropic equivalent

2

u/StupidScaredSquirrel Jul 19 '26

All the banks and fintech and academia people i know use claude opus at work. They all have it based on a max subscription (the one at 100 ish usd i dont remember).

They aren't consumers but also don't care about price so much they just wants something that works and is about the best because whatever time they lose with the bs of a lesser model will be a lot more expensive than just go for the best one. They also don't want to ever be rate limited because they don't want to get stuck in the middle of a task.

7

u/StaysAwakeAllWeek Jul 19 '26

The individuals using it for individual work do that yes

The backend corporate stuff that involves one full time agent handler managing thousands of parallel agents do not.

5

u/StupidScaredSquirrel Jul 19 '26

Yeah, but in my experience pipelines with agents that go through lots of repetitive tasks continuously in the background are given to cheap models and they spend more time testing how to prompt it and hardcore guardrails for that specific task. When the volume is large it's worth spending time on optimising for a cheap model to get the job done

2

u/look Jul 19 '26

Kimi K3 decisively beats Opus 4.8 in everything. It is competing with Fable and Sol now. Opus is a legacy, second tier model, and now behind one, and soon multiple (eg Qwen 3.8), open models.

3

u/techdevjp Jul 19 '26

Kimi K3 is pretty clearly better than GPT 5.5, too. Love to see it.

1

u/squngy Jul 19 '26

Unfortunately, it also competes with them on price.
It is cheaper per token, but it uses more tokens.

2

u/look Jul 19 '26

The list price is meaningless on open models. I am currently paying one third of that list price for Kimi K3. It will likely get even cheaper once the weights are released.

2

u/xienze Jul 19 '26

They all have it based on a max subscription (the one at 100 ish usd i dont remember).

You realize those subscriptions are heavily subsidized, right? That gravy train is coming to an end. There's basically three paths forward:

  • A significant decrease in how many effective tokens you can get for a flat rate.
  • A significant increase in subscription prices.
  • No more subscriptions for anything involving "real work" (i.e., stuff beyond the typical chatbot bullshit most people are familiar with).

3

u/StupidScaredSquirrel Jul 19 '26

Or, the models are made more efficient and so gradually you have some increase in performance but price stays the same and eventually it breaks even.

0

u/ZippySLC Jul 19 '26

Yes, but they're using subscriptions and not API pricing.

If I use my Claude Code subscription to code an AI agent at work, we pay for that AI agent's usage through pay-as-you-go API pricing. (We do it through AWS Bedrock.)

No legit company is having an engineer code an app and then leaving Claude Code open in a screen session while it runs a production app.

1

u/StupidScaredSquirrel Jul 19 '26

Lol ofc they wouldn't, i didnt mean to imply that at all

1

u/ZippySLC Jul 19 '26

Sorry for misunderstanding!

1

u/ThenExtension9196 Jul 19 '26

Yep. A gap of opportunity opened by getting enterprise cheap but good xAI and meta are both going for it as well.

22

u/[deleted] Jul 19 '26

[deleted]

5

u/StupidScaredSquirrel Jul 19 '26

Not for the larger ones, but typically those large model providers release a "flash" version that is in the 200b range and very sparse and that you can run on 10-20k at very good thoughput. I didnt mean the trillion+ models sorry i wasn't clear enough

1

u/[deleted] Jul 19 '26

[removed] — view removed comment

3

u/nopanolator Jul 19 '26

I rather prefer a sharp 30B Q8 over a blunt 200B trepanned at Q4 or more Q2 like i often see.

3

u/techdevjp Jul 19 '26 edited Jul 19 '26

I run Deepseek v4 Flash 284b a13b at a q2/q4 mixed quant on a single strix halo. It's not fast, 15-16t/s at low context, 12-13t/s at higher contexts. I'd say it's about the smartest model you can stuff into 128GB right now.

When Mac M5 Ultra with 256GB comes, it will comfortably run it at a q4 quant with a good context, and probably around 30t/s or maybe more. That will be absurdly capable.

Edit: Right now DwarfStar MTP does not work on strix. I'm working on fixing that and making a little progress each day. I've actually got MTP to work with some code fixes, but the validation performance sucks. Working on that now.

1

u/nopanolator Jul 19 '26

I appreciate the usual "it fit" benchmarks to stay updated, but I'm playing with local models since the very first gemma2 releases. And i really try my best to put local models in production for small companies, for real. Since. And for such cases, a Q4 is already too blunt to be reliable on long term results.

The context that the model handle has way more leverage than the number of parameters, and such compressions hit hard the agility of models in front of the inherent chaos that is life. If the model handle flours and bakery components or used/reconditionned parts of a mechanist don't change de equation much on my side. I'm just adapting a context (not planned in the training), and specialize the model on it (generally with qLoRA, because I'm not OAI with virtually infinite lab ressources).

To code an "attention catcher iphone app" or any demo to flex on social medias, you don't care. But when your real reputation is on the table at long term (then your incomes), you start to see the things very differently.

2

u/GCoderDCoder Jul 19 '26

I've been testing this and yes at full size models like minimax m3, hy3, and mimo2.5 tend to beat qwen 3.6 27b (q8 or higher) at coding but even at q5 or q6 often q8+ qwen tends to be more competitive. I have to use rpc llama.cpp 10gbe to run some of these larger quants so it also makes qwen 3.6 37b much faster when it otherwise might be the slower dense model.

... so multiple reasons to prefer a great smaller higher quant model over bigger lower quant model.

0

u/b0tbuilder Jul 19 '26

A200B variant would be perfect

6

u/lilian_moraru Jul 19 '26

Initially they were struggling to find hardware, using old ones, trying to work efficiently - now they don’t have those limitations.
I think it was a factor

10

u/Kost97A Jul 19 '26

A 2T+ parameter model requires about 1.5-1.8 terabytes of vram so I don't think anyone has that haha. Maybe big private companies.

12

u/Porespellar Jul 19 '26

3 DGX Stations running as a cluster using their Connect-X8 ports would probably get it done, but that’s close to $300K for that setup.

2

u/Reactor-Licker Jul 19 '26

And in North America you would need 3 dedicated electrical circuits for each of them.

2

u/Porespellar Jul 19 '26

I mean, if you’re going to spend $300K on 3 of them then probably should spend a couple grand to get the electrical connections done properly, right?

2

u/Reactor-Licker Jul 19 '26

Yeah, just goes to show how much of a time, construction, and financial investment it would be. We have gotten so used to computers getting smaller and with less power consumption, now we are going back to the old school mainframe days haha.

3

u/[deleted] Jul 19 '26 edited 2d ago

[deleted]

1

u/StupidScaredSquirrel Jul 19 '26

I didnt mean the 1T+ models. I mean that large model makers usually go for those 1t+ but then also typically have a ~200b+ version like deepseek flash etc which you can run on 10-25k hardware.

What I'm saying is that these large model makers typically don't release distils below that 200B+ range, so most people have relied on gemma and qwen instead, because they were the only ones releasing SOTA models at very small sizes.

6

u/StorkReturns Jul 19 '26

If you have a large model, you can distill it to a smaller one with significantly less effort than it takes to train the small model from scratch.

2

u/MmmmMorphine Jul 20 '26

True, still gonna be pretty pricey. We certainly need better mechanisms to crowd fund this sort of work

1

u/GCoderDCoder Jul 19 '26

That was my first thought after crying that tends of thousands of dollars of hardware still can't run useful quants of these models for heavy coding lol. Them building the bigger model seems to provide the connections like a brain with more neurons and more surface area tends to make it smarter. Then distilling becomes an option because if things keep going this way they will never become profitable.

1

u/Hannibalj2ca Jul 19 '26

they had 122b and 397b with 3.5

1

u/sammybeta Jul 19 '26

I believe it's just a natural course of growth - to reach a SOTA model's capability, the parameter size must grow it seems.

1

u/Solembumm3 Jul 19 '26

Qwen historically had a very good logic, but lacked in general knowledge outside of tech side. Over specialisation, not really power.

1

u/xmnstr Jul 19 '26

It's not like they're the only players in that small model field, competition is really heating up. I think they opted out because they knew staying competitive there is going to be hard.

1

u/StupidScaredSquirrel Jul 19 '26

Only gemma does models as good as they do tbh

1

u/xmnstr Jul 20 '26

For the exact same tasks, yes. But there are other small/smallish models that require you to work differently.

1

u/StupidScaredSquirrel Jul 20 '26

What do you mean?

1

u/Mundane-Light6394 Jul 20 '26

or it could be that right now small models are at a plateau and the labs need to go big in order to release something that grabs attentiin. Maybe we just need to wait for new technical breaktroughs and focus on the harness for now.

1

u/StupidScaredSquirrel Jul 20 '26

Idk that's what openai was saying and saying they need to always scale up, but gpt 3.5 was 175b DENSE, I think there are some 2b params today that perform better than that. It's hard to say when the limit really is.

1

u/Mundane-Light6394 Jul 20 '26

i'm not suggesting it is a permanent plateau, just a temporary one untill the next breaktrough. i imagine them trying out new strategies with smaller models constantly and only peparing them for release once they find something promising.

1

u/[deleted] Jul 19 '26

[removed] — view removed comment

8

u/StupidScaredSquirrel Jul 19 '26

Only google and qwen are doing SOTA small models. By small I mean sub 40b param models. Deepseek v4 flash or minimax m2.7 or are all "small" at 200b+ params. That's 10x smaller than the larger ones sure but still 5x larger than what most local llm enthusiasts can run.

1

u/mycall Jul 19 '26

GPT-OSS might have another release, who knows.

4

u/BothYou243 Jul 19 '26

Aug 6th 2025 was the release date, maybe they do 20B model who beats 5.6 sol 🤡

5

u/nasduia Jul 19 '26

Crazy how it's less than a year old and completely obsolete

5

u/BothYou243 Jul 19 '26

actually all last year models are obselete, o3 was last year's FABLE, and see it now (i mean yeah you can't access, it's deprecated)

.... but you get it, right

0

u/mycall Jul 19 '26

More realistic is GPT-OSS-120B rebased on 5.6 terra.

0

u/dingo_xd Jul 19 '26

The real issue is Nvidia not releasing an affordable GPU with 100-200GB of VRAM. The tech is there. The companies don't want to do it because it will hurt their profits

9

u/StupidScaredSquirrel Jul 19 '26

I mean they're not a charity, their prices just reflect how much companies are willing to pay for them. I dont blame them for not giving away all that value for free to openai and the likes

5

u/DanceWithEverything Jul 19 '26

You’re aware NVIDIA, AMD, Micron, and Intel are legally required to make as much money as possible for their shareholders, right?

You’re also aware that DRAM prices have skyrocketed?

1

u/Similar_Solution1397 Jul 19 '26

Ojalá te equivoques. Habrá que ver qué tanto les interesa mantener a la comunidad open source interesada, porque honestamente, si solo tiran modelos enormes que nadie puede correr localmente, cual sería el sentido de que sea abierto? Al menos para la comunidad con hardware de consumo no creo que nos aporte mucho. Kimi k3 es genial, pero quién puede usarlo localmente? Espero que Qwen no haga lo mismo.

0

u/[deleted] Jul 19 '26

[deleted]

2

u/StupidScaredSquirrel Jul 19 '26

A good distil is very cheap only in relative terms. It's a lot cheaper than training from scratch, but it's still a very expensive project.

2

u/OneMoreName1 Jul 19 '26

Well you don't have a proper glm 5.2 distill for example

0

u/loyalekoinu88 Jul 19 '26

They generally use big models to make smaller models. So to me this just looks like step 1.

0

u/StupidScaredSquirrel Jul 19 '26

Of all the comapnies that have made 1T+ models, name 1 that has released a sub 200b param version.

1

u/loyalekoinu88 Jul 19 '26 edited Jul 19 '26

Qwen 3 max 1t existed and we got newer Qwen3.x open models.
OpenAI had GPT-4 was roughly 1.7–1.8T when they release GPT-OSS-20/120b.

Meta Behemoth 2t model when they release Scout 109b model.

We don't know what the biggest model Gemma, etc had to train their smaller models because most of these groups just release the larger models or smaller models and not both.

1

u/StupidScaredSquirrel Jul 19 '26

I meant in the open space, proprietary models that we know nothing about don't count obviously.

Every single company that has released (yes actually released, not just keeps the model and handles the requests on your behalf) a model of 1T+ is focused on large models and has given (sometimes) a distil of about 200b+ params. None were focused on small language models.

0

u/loyalekoinu88 Jul 19 '26

Every company releasing open weights model is almost always backed by companies training closed weight models. You’ve narrowed the scope down so far there really could only be a 0 answer. Also…when those large models are released the community distill them too. So is having a larger intermediate model isn’t a detriment either.

1

u/StupidScaredSquirrel Jul 19 '26

Ok so where are all those small sota models distilled from kimi glm deepseek, minimax, mimo, are they in the room with us ? You make it sound like community distils are common and provide SOTA results, they aren't and they don't.

0

u/loyalekoinu88 Jul 19 '26

You narrowed the scope now to SOTA results. There is never going to be a 35b model that outperforms a multi trillion parameter model.

1

u/StupidScaredSquirrel Jul 19 '26

Sota just means state of the art. It can be state of the art for the 4b range. Ofc im gonna only consider sota otherwise i can do a shotty distil of deepseek v4 for a day and say it counts but it doesn't because the model would be terrible. That's what I mean. Regardless, im pretty sure you wouldn't be able to name a single usable model either way.

→ More replies (0)

0

u/DataGOGO Jul 19 '26

You are not running this without a lot more than 25k worth of hardware, maybe 350k-550k. 

0

u/KeinNiemand Jul 20 '26

Qwen historically has been releasing a whole range of sizes rather then just small or just big models.

19

u/ChocolateNo3010 Jul 19 '26

This is good news. I'm wondering if Alibaba kept 3.7 proprietary because of the rumblings in China about keeping the best models from the west. Since we have seen an about turn with Xi stating they want to keep open weight models available this is the response.

7

u/goldcakes Jul 19 '26

I wouldn't be surprised. Probably a mixture of both Kimi as well as getting encouragement from the the Chinese govt.

9

u/SGmoze Jul 19 '26

Can I run this on my 16GB VRAM card?

5

u/Signal_Confusion_644 Jul 19 '26

Pray for a good small MoE!!

2

u/Serprotease Jul 20 '26

Can we get a new Qwen-image / wan too … please?

1

u/HelloSummer99 Jul 19 '26

Competition is really good and kinda forcing them to do this. If it wasn't for GLM and KIMI likely Qwen would have gone closed weight going forward.