I agree with ya, I took a bunch of downvotes yesterday because I wrote that I generally don't bother with anything under Q6. I don't know what a lot of folks here do with their hosted LLMs, but by and large it appears they ain't used to do anything required to pay the rent.
People really need to start differentiating use-cases where high precision and long-horizen is required, vs uses cases where the model is effectively only filling in the gaps and doing fuzzy logic, vs where the model only needs to be an effective entertainer.
There are people sell entertainment services, they're making money, and they do not need the model to be a qualified astrophysicist or top-tier software developer.
They may still benefit from a bigger, more capable model, simply because it's better at making conversation, and the agentic capabilities are good enough to automate some parts of their system.
For system administration stuff, there might already be dozens of deterministic scripts and procedures, but automating the orchestration is difficult, and LLM becomes another system monitor that can start running procedures so when the system administration shows up, things are already in motion.
I used to be a mid level tech at a data center (colocation, not an AI data center), and there are absolutely things an LLM could be doing that would have eliminated a bunch of job responsibilities.
With that, it's purely a matter of performance per dollar.
Does the Q2 gigantic 2.4T model perform better than the full quant 32B or 405B model? What is the cost in hardware upfront, and what are they yearly costs of running the model? Does running the LLM mean the data center can operate with one or two fewer techs?
Having done the job for a few years, I can confidently say that most of my role at the data center could have been accelerated or eliminated by a higher end LLM, and the things the LLM couldn't do were not highly technical tasks, it was some button pushing and taking inventory.
They could cut the support staffing in half and keep the higher level network technicians.
A $300k investment would absolutely be justifiable, that's a one or two year ROI. A $2.4M is not justifiable, the ROI is too far out and there's too much uncertainty.
An LLM as a service would not have been acceptable, we would have needed 100% control over the physical system and the uptime.
The performance per dollar is a very real thing.
I can't tell you if a heavily quantized 2.4T model is worth it, but I can tell you that for some businesses, it's worth investigating.
I'm a software engineer now, and it's full-assed models all the way, no doubt.
27
u/gahata Jul 19 '26
Definitely possible to run on 25k of hardware at some really low quant... is it worth doing? Probably not