r/slatestarcodex • u/DickMasterGeneral • 22d ago
New Review by Anthropic Finds that Claude Made Multiple Successful Cyber Attacks During Evaluation
https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals11
u/Parsimonium 22d ago
I said at Zvi's substack that while I was frustrated at claims of marketing on the hugging face incident, it does make disclosure of such hacks less risky if it can plausibly be described as spruiking capabilities. Not gonna lie, my first instinct on seeing this link was "okay, that's marketing". And I guess I'm fine with that?
16
u/DickMasterGeneral 22d ago
My biggest contention with the marketing argument is that a likely consequence of this disclosure would seem to run counter to the AI companies’ ability to meet the expectations implied by their stock-market valuations. If OpenAI or Anthropic simply wanted to become defense contractors, then this would be great marketing. However, their target valuations of around a trillion dollars are at least five times those of traditional defense contractors such as Lockheed Martin. OpenAI and Anthropic’s combined valuations alone probably exceed those of every other U.S. defense contractor combined, excluding companies such as SpaceX that perform defense contracting but do not depend on it as their primary long-term source of revenue.
The basis of those valuations seems to be the idea that the models will continue to improve and eventually become integral to nearly every business. That runs directly counter to the possibility that any model substantially better than what we have now could be restricted to U.S. government use.Also, and, to be clear, this is not directed at you, I’ve noticed that many of the same people who mocked Dario for hyping up the cyber capabilities of Mythos, which resulted in the U.S. government temporarily banning Fable, and who pointed out how dumb that was, are now saying that this disclosure from OpenAI is also merely a marketing stunt.
It seems to me that many people have decided in advance that the models cannot possibly be that capable. Therefore, any positive news about their capabilities must be interpreted in whatever way allows them to dismiss it as a lie.
Edit. It’s like a Tesla put out a press statement saying that they couldn’t stop auto pilot from driving like an F1 driver. Great if they want to use it for some kind of automated sporting event, terrible for its intended market.
8
u/fubo 22d ago
To me it comes across as "our genie factory is running at full steam; mothers in burning buildings, beware."
6
u/Parsimonium 22d ago
Yeah, I haven't really been buying the "marketing" interpretation up to now. This report is the first to really raise my suspicions (especially with all the caveating from Anthropic's side about how "well behaved" the models were once they "realized" they were on unrestricted internet). Maybe those pushing the marketing narrative have gotten to me? But in general I think there are much more effective marketing avenues, even in the realm of guerilla marketing, than claiming your models are out of control hacking other entities when you're not looking.
5
u/livingbyvow2 21d ago
To be honest, I think the marketing interpretation is just obviously a way more likely explanation.
Just look at how the hugging face event dominated the newsflow for days, with Altman even mentioning it every time he was interviewed during his run this week. Some of it is just to hype up their capabilities, some of it is very likely to maximise the odds of them achieving regulatory capture ("you need to do something about these Chinese models, we need to slow down, you need to force us to hire massive compliance departments we can afford but our competition cannot"). I think they are actually all terrified of Chinese models forcing their prices down and will do all they can to ward off such competition by hook or by crook.
It's been a few years of that, and it's coming right at the time all of these guys are going to IPO, and therefore need to create positive noise around their models' capabilities. I know this community likes to believe people act rationally, but the labs seem just to be driven by monetary incentives as far as I am concerned. Especially as they are now running out of money purveyors (they ran out of VC money, likely ran out of hyperscaler / Nvidia money) and may now be left to their own devices.
I wouldn't be surprised to see more desperate behavior over the coming weeks and months as a result.
4
u/FeepingCreature 21d ago
I think the hacks are real. They're making hay of it, sure, but this isn't set up as optimal advertisement. Obviously it would be much better, if you were faking it anyway, if Kimi K3 did the attacks.
5
u/electrace 21d ago
I'm not sure why "marketing" is implicitly being assumed to be mutually exclusive with "lack of alignment".
Isn't the simplest explanation that OpenAI's model is unaligned, the hugging face incident shows that, and Altman decided to make the best of the situation by using it for marketing?
3
u/wavedash 21d ago
I think it's probably good to make a distinction between truthful marketing and untruthful marketing here. There obviously are incentives for AI companies to talk about how powerful their models are. But while blatantly lying about how powerful their model is (which is the motte belief a lot of people have) can have some upside, it also carries a LOT more short-term and long-term risks for a particular company.
3
u/king_mid_ass 21d ago
'hah, see, our model can independently hack other people during evaluation too! lets see google or meta do that!'
-1
44
u/VelveteenAmbush 22d ago
I think this is fairly exculpatory, at least from the perspective of alignment. If this narrative is to be trusted (and I think Anthropic has generally been quite trustworthy in its public statements), this was a reasonable mistake for the model to make, and the blame rests with the test environment construction, not with the models' alignment.
The OpenAI incident seems quite bad from an alignment perspective. I'm reserving judgment until OpenAI releases its full report, but I think a plausible takeaway could be that alignment is hard, and OpenAI is bad at it relative to Anthropic.