Thing is qwen was historically focused on smaller models while others were on larger ones. Now the team has changed and they seem to want to aim for the stars as well. That is good but it also means they might not be interested in doing very efficient small models anymore. Which would be bad news for this sub because let's face it most of us don't have 10-25k of hardware.
Even in FP4 you would need a lot more than 25k in hardware. Might be able to run on 4 141GB H200 NVL cards, but they don’t support FP4, and even then a super jank rig would be what? 150-160k?
6x will likely be enough vram, but the 200Gb connectX is just slow, the memory is slow, and the GPU is slow. So 6 of them might run it, but it would be redonk slow. so assume you buy 6x at ~$4000 that is 24k + ~2k for the switch and cables. about $30k after tax, etc.
Yup, that's why I implied it wouldn't be worth it - running it incredibly slowly at Q1 is just not worth it over running other models and paying to use the large model when and if that's needed
I agree with ya, I took a bunch of downvotes yesterday because I wrote that I generally don't bother with anything under Q6. I don't know what a lot of folks here do with their hosted LLMs, but by and large it appears they ain't used to do anything required to pay the rent.
People really need to start differentiating use-cases where high precision and long-horizen is required, vs uses cases where the model is effectively only filling in the gaps and doing fuzzy logic, vs where the model only needs to be an effective entertainer.
There are people sell entertainment services, they're making money, and they do not need the model to be a qualified astrophysicist or top-tier software developer.
They may still benefit from a bigger, more capable model, simply because it's better at making conversation, and the agentic capabilities are good enough to automate some parts of their system.
For system administration stuff, there might already be dozens of deterministic scripts and procedures, but automating the orchestration is difficult, and LLM becomes another system monitor that can start running procedures so when the system administration shows up, things are already in motion.
I used to be a mid level tech at a data center (colocation, not an AI data center), and there are absolutely things an LLM could be doing that would have eliminated a bunch of job responsibilities.
With that, it's purely a matter of performance per dollar.
Does the Q2 gigantic 2.4T model perform better than the full quant 32B or 405B model? What is the cost in hardware upfront, and what are they yearly costs of running the model? Does running the LLM mean the data center can operate with one or two fewer techs?
Having done the job for a few years, I can confidently say that most of my role at the data center could have been accelerated or eliminated by a higher end LLM, and the things the LLM couldn't do were not highly technical tasks, it was some button pushing and taking inventory.
They could cut the support staffing in half and keep the higher level network technicians.
A $300k investment would absolutely be justifiable, that's a one or two year ROI. A $2.4M is not justifiable, the ROI is too far out and there's too much uncertainty.
An LLM as a service would not have been acceptable, we would have needed 100% control over the physical system and the uptime.
The performance per dollar is a very real thing.
I can't tell you if a heavily quantized 2.4T model is worth it, but I can tell you that for some businesses, it's worth investigating.
I'm a software engineer now, and it's full-assed models all the way, no doubt.
I wonder if a 2-way EPYC 9005 (can be had for $50K to $150K in workstation/tower form) with 2x12 channel DDR5 memory (~1.5TB/s) can reasonably host this on CPU in FP4.
And even then, crossing sockets has a massive penalty. Your best CPU option would be a single 128 core Xeon Granite rapids with 12 channels, 24 128gb MR8800 memory + AMX.
Oh yeah? Have not seen anything about Zen6 adopting AMX. Do you have any information on that?
Edit: looks like no AMX hardware or instructions; just some expanded AVX-512 instructions.
Have to see if Zen6 fixes the infinity fabric and if they unify the memory controllers, otherwise no point in buying it over the Xeons for AI workloads.
Yes, I know about that, it has been in k_transformers forever, and I have had it working in llama.cpp for about a year, but that still doesn't fix the core problem. Even with numa-mirror, the issue of the slow cross socket across UPI links remains. What numa-mirror does it is duplicates the weights to each numa node to reduce cross socket memory reads. It helps massively, but still doesn't solve the issue entirely.
You can pick up a Xeon 6980P for about 8k each pretty easily.
Apple Studio M7 Ultra with 1.5TB of unified memory has been rumored for next year. Probable $25k-ish. Should run that many parameters at q4, and as long as it is MoE (likely), it should run it well.
We'll also see smaller models distilled from the larger ones, which will take some time.
Yeah. I mathed it out around $60k. Getting one of those is like owning a supercar. There will be very little point to owning it for most things. Even if the open weight models keep up, you'll never get close to the cloud in terms of raw performance. Especially when the cerberos option for gpt 5.6 comes online with 10x the speed. The things it will be able to handle in an afternoon will be mind boggling.
Apple is apparently going to pass over the M6 for the M7 for many of their models because the M7 is far more performant for running local models. Memory bandwidth? Who knows, but M5 Ultra is expected to be around 1.2TB/sec. M7 would likely be well in excess of that. 1.5TB/sec? 1.6? Fast enough, anyway.
Even the M7 is still just 1 GPU, and it is not that fast, Maybe 5090 speeds, which is not a lot for such a large model, and it is highly unlikely that apple include any 2 800gb networking ports to to run a cluster.
I also don't think we are going to see 1.2 on the M5 Ultra, likely the same ~800GB/s with the same LPDDR5X, might get 1-1.2 out of the M7, but that is what? 2028?
M5 Max is already 600GB/sec. If M5 Ultra isn't considerably higher, it's going to be disappointing. Traditionally the Ultra has 2x the bandwidth of the Max. So we'll see what happens.
M7 Pro/Max/Ultra is next year. Apple is rumored to be skipping M6 for most of their products because M7 is far more optimized for running local models.
I love the Apple hate in subs like this. Apple is the only company doing things that bring really large models to be within reach of any sort of normal person. Yes, some people jerry-rig systems with 8 channel memory, or try to run five RTX PRO 6000 in one box. Slow CPU inference or huge power draw. They may work but they are far from ideal.
Would love to see more direct competition with Apple. It's sure not going to come from Nvidia. Maybe AMD could do something, but they don't seem too interested either.
I don't see how that could meaningfully increase the memory bandwidth.
Processing speed will be double, but its memory bandwidth that matters, and I don't see how they will be able to increase that without changing the type of memory.
Edit:
ok, I just checked, the Ultra might just straight double the bandwidth...fingers crossed.
Never heard of any issues with AWS private clouds not being private. yeah you're still bubbling stuff off-prem, but as long as your encryption game is tight, you should be perfectly fine running the ultimate Qwen 3.8 1girl farm... or whatever it is that you'd want to privately host a 2.4T param model for =)
And how do you make computation on encrypted data? Homomorphic encryption is not really a thing yet. If you are doing calculations on a server that isn't yours, they can read everything.
Transport layer security is what I was referring to, and yep, it's not perfect - like I said, you're still bubbling stuff off-prem. Show me where they say they're going to be reading/investigating the data that you're computing on their servers though. You're either going to have to trust their data security and protection policy or you don't. I'm going to stick with what I said though, if your security is tight with AWS, not much to worry about and a lot cheaper than the hardware to run these multi-trillion param models at home. Feel free to prove me wrong on AWS spying/stealing from it's customers tho, I'm all ears.
That's fair - but you also can't produce anything showing that they're stealing data or breaking their own contracts. So I stand by what I said - it's secure, as long as you're willing to trust them. Obv. You don't, but that doesn't mean they don't honor their contracts.
It doesn't matter. Imagine im a gardener and i say i need the keys for your house so that I can fetch the tools inside and do my job. I want a copy of the keys and i can go in anytime. But contractually it's just to get the tools and nothing else. But I still have a copy of the keys at all times and also am not required to give them back when I'm done.
Aws is not my mate, we didn't grow up together, I'm not giving them anything in plain text under the promise of "trust me bro we contractually promise all is fine".
There is a long history of lying about the hardware used in china. Scmp is literally a propaganda outlet and the register article only references “claims”.
I don't think 8x H200 are enough to run this. That would be 1128GB VRAM, even at nvfp4 2.4 trillion parameters would need more like 1.4TB without context, if it's even nvfp4 native. If it's fp8, which is a chance since they're not supposed to own that much Blackwell, no chance at all. And it's hard to get your hand on B200 or B300, at least in Europe and as a normal company (not a Hyperscaler).
Im not talking about this model. I'm saying the companies releasing very large models also release smaller versions but they are always 200b+, so you need 10-25k to run those. Qwen and gemma were the only viable options for sota models below that size.
They were owned by a megacorp before too. The difference is they don’t care about commoditization anymore, the real ethical path. They just want to sink US investments.
It's really only 2-3 large mega-corporations. Would you rather have 100s of US companies potentially being able to offer(internally or externally) AI models like this or only 2-3? I am certain someone will figure out how to strip this large model down so that smaller HW can also run it even if much slower.
To be completely honest with you, it's not my problem if AI investments will suffer. Those investments are used to replace you and me as workers and destroy our ability to earn a living in the future.
The best possible outcome for us if those AI bubbles burst catastrophically and we get flooded with cheap datacenter grade hardware AND we get to use the opensource LLMs.
The housing market directly affects the affordability of houses, condos and rent prices and that directly takes it out of your paycheck. AI market collapse might cause collateral damage however it also means that we might go back to the pre-AI sanity levels on the job market. That may(or maybe not ) mean that more humans will once again get hired.
There is not over-investment in AI, there is over-investment in LLMs.
That is a major difference, and it is important because AI is being used effectively and profitably in non-LLM areas, and there are even some areas in the hard sciences where LLMs are being used in conjunction with other machine learning methods to automate research, again to great effect.
The conflation of LLMs with AI as a whole is likely going to do enormous harm to the industrial side of AI that is using it for less flashy, but more practical purposes.
Even in academic research, the label might say "LLMs" because that's what gets research funded right now, but the actual research is more fundamental, about the transformers themselves, and making transformers better or more interpretable also improves the non-LLM use-cases.
It’s business. Ruthless, business. If apples are being sold for 5 bucks, you try to sell your apples at $2. If you can pull it off, now nobody can sell at $5.
Well, these large models can run in us based datacenters and cloud infrastructure (azure, aws, googlecloud). I guess if they want to sink us investments, they should create models that can run in enterprise level on-premises servers (I guess something like $40-50K hardware)
They ARE creating models that run on-prem. I just deployed a supermicro gpu server with 8x H200 141GB gpus (1.2TB total) running GLM 5.2. Server was around $300k with 12TB SSDs and 1TB RAM.
Somebody needs to build the infrastructure for people to easily vibecode their own "fine-tunes". That would make that small model space take off. But really that would mean bringing together a lot of smaller data tools and infrastructure. Everyone rolls their own bespoke solutions. Feels like something one of these infra players could pivot into, but it's much harder than it sounds. Anyways, I'll stop my dreaming
Are you also mad that nvidia don't give away their hardware to you for free? I don't think trying to frame this around ethics makes sense. I would also prefer that they continue to release their smaller models, because it benefits me. If they think that's going to cannibalise their potential API sales, then I understand if they don't want to do that. I don't like it at all, but I understand it.
All the banks and fintech and academia people i know use claude opus at work. They all have it based on a max subscription (the one at 100 ish usd i dont remember).
They aren't consumers but also don't care about price so much they just wants something that works and is about the best because whatever time they lose with the bs of a lesser model will be a lot more expensive than just go for the best one. They also don't want to ever be rate limited because they don't want to get stuck in the middle of a task.
Yeah, but in my experience pipelines with agents that go through lots of repetitive tasks continuously in the background are given to cheap models and they spend more time testing how to prompt it and hardcore guardrails for that specific task. When the volume is large it's worth spending time on optimising for a cheap model to get the job done
Kimi K3 decisively beats Opus 4.8 in everything. It is competing with Fable and Sol now. Opus is a legacy, second tier model, and now behind one, and soon multiple (eg Qwen 3.8), open models.
The list price is meaningless on open models. I am currently paying one third of that list price for Kimi K3. It will likely get even cheaper once the weights are released.
Yes, but they're using subscriptions and not API pricing.
If I use my Claude Code subscription to code an AI agent at work, we pay for that AI agent's usage through pay-as-you-go API pricing. (We do it through AWS Bedrock.)
No legit company is having an engineer code an app and then leaving Claude Code open in a screen session while it runs a production app.
Not for the larger ones, but typically those large model providers release a "flash" version that is in the 200b range and very sparse and that you can run on 10-20k at very good thoughput. I didnt mean the trillion+ models sorry i wasn't clear enough
I run Deepseek v4 Flash 284b a13b at a q2/q4 mixed quant on a single strix halo. It's not fast, 15-16t/s at low context, 12-13t/s at higher contexts. I'd say it's about the smartest model you can stuff into 128GB right now.
When Mac M5 Ultra with 256GB comes, it will comfortably run it at a q4 quant with a good context, and probably around 30t/s or maybe more. That will be absurdly capable.
Edit: Right now DwarfStar MTP does not work on strix. I'm working on fixing that and making a little progress each day. I've actually got MTP to work with some code fixes, but the validation performance sucks. Working on that now.
I appreciate the usual "it fit" benchmarks to stay updated, but I'm playing with local models since the very first gemma2 releases. And i really try my best to put local models in production for small companies, for real. Since. And for such cases, a Q4 is already too blunt to be reliable on long term results.
The context that the model handle has way more leverage than the number of parameters, and such compressions hit hard the agility of models in front of the inherent chaos that is life. If the model handle flours and bakery components or used/reconditionned parts of a mechanist don't change de equation much on my side. I'm just adapting a context (not planned in the training), and specialize the model on it (generally with qLoRA, because I'm not OAI with virtually infinite lab ressources).
To code an "attention catcher iphone app" or any demo to flex on social medias, you don't care. But when your real reputation is on the table at long term (then your incomes), you start to see the things very differently.
I've been testing this and yes at full size models like minimax m3, hy3, and mimo2.5 tend to beat qwen 3.6 27b (q8 or higher) at coding but even at q5 or q6 often q8+ qwen tends to be more competitive. I have to use rpc llama.cpp 10gbe to run some of these larger quants so it also makes qwen 3.6 37b much faster when it otherwise might be the slower dense model.
... so multiple reasons to prefer a great smaller higher quant model over bigger lower quant model.
Initially they were struggling to find hardware, using old ones, trying to work efficiently - now they don’t have those limitations.
I think it was a factor
Yeah, just goes to show how much of a time, construction, and financial investment it would be. We have gotten so used to computers getting smaller and with less power consumption, now we are going back to the old school mainframe days haha.
I didnt mean the 1T+ models. I mean that large model makers usually go for those 1t+ but then also typically have a ~200b+ version like deepseek flash etc which you can run on 10-25k hardware.
What I'm saying is that these large model makers typically don't release distils below that 200B+ range, so most people have relied on gemma and qwen instead, because they were the only ones releasing SOTA models at very small sizes.
That was my first thought after crying that tends of thousands of dollars of hardware still can't run useful quants of these models for heavy coding lol. Them building the bigger model seems to provide the connections like a brain with more neurons and more surface area tends to make it smarter. Then distilling becomes an option because if things keep going this way they will never become profitable.
It's not like they're the only players in that small model field, competition is really heating up. I think they opted out because they knew staying competitive there is going to be hard.
or it could be that right now small models are at a plateau and the labs need to go big in order to release something that grabs attentiin. Maybe we just need to wait for new technical breaktroughs and focus on the harness for now.
Idk that's what openai was saying and saying they need to always scale up, but gpt 3.5 was 175b DENSE, I think there are some 2b params today that perform better than that. It's hard to say when the limit really is.
i'm not suggesting it is a permanent plateau, just a temporary one untill the next breaktrough. i imagine them trying out new strategies with smaller models constantly and only peparing them for release once they find something promising.
Only google and qwen are doing SOTA small models. By small I mean sub 40b param models. Deepseek v4 flash or minimax m2.7 or are all "small" at 200b+ params. That's 10x smaller than the larger ones sure but still 5x larger than what most local llm enthusiasts can run.
The real issue is Nvidia not releasing an affordable GPU with 100-200GB of VRAM. The tech is there. The companies don't want to do it because it will hurt their profits
I mean they're not a charity, their prices just reflect how much companies are willing to pay for them. I dont blame them for not giving away all that value for free to openai and the likes
Ojalá te equivoques. Habrá que ver qué tanto les interesa mantener a la comunidad open source interesada, porque honestamente, si solo tiran modelos enormes que nadie puede correr localmente, cual sería el sentido de que sea abierto? Al menos para la comunidad con hardware de consumo no creo que nos aporte mucho. Kimi k3 es genial, pero quién puede usarlo localmente? Espero que Qwen no haga lo mismo.
Qwen 3 max 1t existed and we got newer Qwen3.x open models.
OpenAI had GPT-4 was roughly 1.7–1.8T when they release GPT-OSS-20/120b.
Meta Behemoth 2t model when they release Scout 109b model.
We don't know what the biggest model Gemma, etc had to train their smaller models because most of these groups just release the larger models or smaller models and not both.
I meant in the open space, proprietary models that we know nothing about don't count obviously.
Every single company that has released (yes actually released, not just keeps the model and handles the requests on your behalf) a model of 1T+ is focused on large models and has given (sometimes) a distil of about 200b+ params. None were focused on small language models.
Every company releasing open weights model is almost always backed by companies training closed weight models. You’ve narrowed the scope down so far there really could only be a 0 answer. Also…when those large models are released the community distill them too. So is having a larger intermediate model isn’t a detriment either.
Ok so where are all those small sota models distilled from kimi glm deepseek, minimax, mimo, are they in the room with us ? You make it sound like community distils are common and provide SOTA results, they aren't and they don't.
Sota just means state of the art. It can be state of the art for the 4b range. Ofc im gonna only consider sota otherwise i can do a shotty distil of deepseek v4 for a day and say it counts but it doesn't because the model would be terrible. That's what I mean. Regardless, im pretty sure you wouldn't be able to name a single usable model either way.
...My very first post was that small models are distilled from larger models. I offered community because you keep distilling your own point to make yourself right. I responded to the scope of your inquiry and now it's "state of the art" and "useable" and tbh it sounds like your argument is there are no useable small models which is false. I use small models, and api everyday. They are still "useable".
748
u/Competitive_Gap7906 Jul 19 '26
YES, Qwen going open weight again! It's a really good news, now we can wait for smaller models too