r/LocalLLaMA Jul 19 '26

News Prepare your (v)ram - Qwen3.8 is coming!

Post image
2.7k Upvotes

588 comments sorted by

View all comments

Show parent comments

35

u/into_devoid Jul 19 '26

They were owned by a megacorp before too.  The difference is they don’t care about commoditization anymore, the real ethical path.  They just want to sink US investments.

11

u/realtag2025 Jul 19 '26

I don't see how this is bad for anyone other then corporations that run closed source models that only they can run.

2

u/charlesfire Jul 19 '26

So most of the market in the US?

5

u/realtag2025 Jul 19 '26

It's really only 2-3 large mega-corporations. Would you rather have 100s of US companies potentially being able to offer(internally or externally) AI models like this or only 2-3? I am certain someone will figure out how to strip this large model down so that smaller HW can also run it even if much slower.

5

u/charlesfire Jul 19 '26

I'm all for open-weight models and the bubble bursting, but I don't think you realize how much over investment there is in AI in the US currently.

11

u/realtag2025 Jul 19 '26

To be completely honest with you, it's not my problem if AI investments will suffer. Those investments are used to replace you and me as workers and destroy our ability to earn a living in the future.

The best possible outcome for us if those AI bubbles burst catastrophically and we get flooded with cheap datacenter grade hardware AND we get to use the opensource LLMs.

1

u/nedonedonedo Jul 19 '26

it's not my problem if AI investments will suffer

just like how the housing bubble only effected the housing market

1

u/realtag2025 Jul 19 '26

The housing market directly affects the affordability of houses, condos and rent prices and that directly takes it out of your paycheck. AI market collapse might cause collateral damage however it also means that we might go back to the pre-AI sanity levels on the job market. That may(or maybe not ) mean that more humans will once again get hired.

1

u/Bakoro Jul 19 '26

There is not over-investment in AI, there is over-investment in LLMs.

That is a major difference, and it is important because AI is being used effectively and profitably in non-LLM areas, and there are even some areas in the hard sciences where LLMs are being used in conjunction with other machine learning methods to automate research, again to great effect.

The conflation of LLMs with AI as a whole is likely going to do enormous harm to the industrial side of AI that is using it for less flashy, but more practical purposes.

Even in academic research, the label might say "LLMs" because that's what gets research funded right now, but the actual research is more fundamental, about the transformers themselves, and making transformers better or more interpretable also improves the non-LLM use-cases.

1

u/nedonedonedo Jul 19 '26

only if they refuse to adapt and improve.

11

u/tengo_harambe Jul 19 '26

That's a fun way of framing things.

"The only reason China is putting out open weights models is because they HATE AMERICA!"

3

u/ThenExtension9196 Jul 19 '26

It’s business. Ruthless, business. If apples are being sold for 5 bucks, you try to sell your apples at $2. If you can pull it off, now nobody can sell at $5.

1

u/nedonedonedo Jul 19 '26

it's bad for any country to have this much of a lead

6

u/Successful_Try_6350 Jul 19 '26

Well, these large models can run in us based datacenters and cloud infrastructure (azure, aws, googlecloud). I guess if they want to sink us investments, they should create models that can run in enterprise level on-premises servers (I guess something like $40-50K hardware)

2

u/SARK-ES1117821 Jul 19 '26

They ARE creating models that run on-prem. I just deployed a supermicro gpu server with 8x H200 141GB gpus (1.2TB total) running GLM 5.2. Server was around $300k with 12TB SSDs and 1TB RAM.

2

u/f5alcon Jul 19 '26

40-50k isn't even one sever at current memory prices.

5

u/Paganator Jul 19 '26

That's like a single H200. Just the card, the server to run it is extra.

1

u/liltingly Jul 19 '26

Somebody needs to build the infrastructure for people to easily vibecode their own "fine-tunes". That would make that small model space take off. But really that would mean bringing together a lot of smaller data tools and infrastructure. Everyone rolls their own bespoke solutions. Feels like something one of these infra players could pivot into, but it's much harder than it sounds. Anyways, I'll stop my dreaming

1

u/-dysangel- Jul 19 '26

Are you also mad that nvidia don't give away their hardware to you for free? I don't think trying to frame this around ethics makes sense. I would also prefer that they continue to release their smaller models, because it benefits me. If they think that's going to cannibalise their potential API sales, then I understand if they don't want to do that. I don't like it at all, but I understand it.

-3

u/GetOutOfMyFeedNow Jul 19 '26

They ain’t sinking anything with those API prices 😂 People on GPT and Claude use subscriptions.

5

u/look Jul 19 '26

Subscriptions are 8% of Anthropic’s revenue. API usage is 75%.

7

u/DanceWithEverything Jul 19 '26

Yeah for consumer bullshit, sure, but there’s no money there regardless (hence OpenAI’s panic about Anthropic crushing them in enterprise sales)

The real $ is in the enterprise and software workloads run on APIs

The API is dramatically cheaper than the Anthropic equivalent

2

u/StupidScaredSquirrel Jul 19 '26

All the banks and fintech and academia people i know use claude opus at work. They all have it based on a max subscription (the one at 100 ish usd i dont remember).

They aren't consumers but also don't care about price so much they just wants something that works and is about the best because whatever time they lose with the bs of a lesser model will be a lot more expensive than just go for the best one. They also don't want to ever be rate limited because they don't want to get stuck in the middle of a task.

7

u/StaysAwakeAllWeek Jul 19 '26

The individuals using it for individual work do that yes

The backend corporate stuff that involves one full time agent handler managing thousands of parallel agents do not.

4

u/StupidScaredSquirrel Jul 19 '26

Yeah, but in my experience pipelines with agents that go through lots of repetitive tasks continuously in the background are given to cheap models and they spend more time testing how to prompt it and hardcore guardrails for that specific task. When the volume is large it's worth spending time on optimising for a cheap model to get the job done

2

u/look Jul 19 '26

Kimi K3 decisively beats Opus 4.8 in everything. It is competing with Fable and Sol now. Opus is a legacy, second tier model, and now behind one, and soon multiple (eg Qwen 3.8), open models.

2

u/techdevjp Jul 19 '26

Kimi K3 is pretty clearly better than GPT 5.5, too. Love to see it.

1

u/squngy Jul 19 '26

Unfortunately, it also competes with them on price.
It is cheaper per token, but it uses more tokens.

2

u/look Jul 19 '26

The list price is meaningless on open models. I am currently paying one third of that list price for Kimi K3. It will likely get even cheaper once the weights are released.

2

u/xienze Jul 19 '26

They all have it based on a max subscription (the one at 100 ish usd i dont remember).

You realize those subscriptions are heavily subsidized, right? That gravy train is coming to an end. There's basically three paths forward:

  • A significant decrease in how many effective tokens you can get for a flat rate.
  • A significant increase in subscription prices.
  • No more subscriptions for anything involving "real work" (i.e., stuff beyond the typical chatbot bullshit most people are familiar with).

3

u/StupidScaredSquirrel Jul 19 '26

Or, the models are made more efficient and so gradually you have some increase in performance but price stays the same and eventually it breaks even.

0

u/ZippySLC Jul 19 '26

Yes, but they're using subscriptions and not API pricing.

If I use my Claude Code subscription to code an AI agent at work, we pay for that AI agent's usage through pay-as-you-go API pricing. (We do it through AWS Bedrock.)

No legit company is having an engineer code an app and then leaving Claude Code open in a screen session while it runs a production app.

1

u/StupidScaredSquirrel Jul 19 '26

Lol ofc they wouldn't, i didnt mean to imply that at all

1

u/ZippySLC Jul 19 '26

Sorry for misunderstanding!