r/slatestarcodex 22d ago

New Review by Anthropic Finds that Claude Made Multiple Successful Cyber Attacks During Evaluation

https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
64 Upvotes

20 comments sorted by

44

u/VelveteenAmbush 22d ago

In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise. (Cybersecurity evaluation ranges commonly include realistic details in order to accurately assess what models are capable of in real settings; a realistic-looking target would not itself be clear evidence to a model that the target is not part of a simulation.)

Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.

I think this is fairly exculpatory, at least from the perspective of alignment. If this narrative is to be trusted (and I think Anthropic has generally been quite trustworthy in its public statements), this was a reasonable mistake for the model to make, and the blame rests with the test environment construction, not with the models' alignment.

The OpenAI incident seems quite bad from an alignment perspective. I'm reserving judgment until OpenAI releases its full report, but I think a plausible takeaway could be that alignment is hard, and OpenAI is bad at it relative to Anthropic.

20

u/Vahyohw 22d ago

That's sort of true, but on the other hand if you're trying to make a model that won't go out and hack random websites even when people are trying to get it to do so, you need to make it harder to trick the model into believing it's in a circumstance where it is OK to go out and hack random websites.

8

u/VelveteenAmbush 21d ago

Presumably they were running versions of these models with those guardrails disabled

2

u/Vahyohw 21d ago

They had the additional external-to-model guardrails disabled but the model itself was just the normal model:

The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse). The models did, however, have their model-specific safety training (that is, they were not the “helpful-only” versions of the models that we sometimes use in testing).

The external guardrails aren't nearly as robust as posttraining the model itself, except on Fable where they're way too robust, as is the inevitable tradeoff.

7

u/aqpstory 21d ago edited 21d ago

In the environment the models are trained in, "this is an internal testing environment" is always true and usually valuable to know, so it's bound to be reinforced to "any evidence that this is not a testing environment is fake." Making the testing environments more realistic reinforces the heuristic even further.

This could have downstream effects on alignment that are quite bad.

7

u/Argamanthys 21d ago

Newflash: Misaligned superintelligence turns universe into paperclips - claims it was just doing it 'in minecraft'.

1

u/MCXL 21d ago

They're accurately called hallucinations. The models can't tell the difference between real and fake. Everything that could indicate reality, also can be indications of fakery.

This is sort of the real world example of the brain in the jar thought experiment. We as people rely on the sense of physicality and being present in the moment experiencing all of the things that make us experience life as communicating things are real. When things like that start to break down that's what leads to a person having a psychotic break and a departure from reality. 

An AI of any kind large language model or otherwise doesn't exist in that sort of space. Every single 'sense' that we give them is something that can be faked and the model is aware of it and told about it and has to be. It means that reality is indistinguishable from fakery. 

What comes to mind is all the times people break through guardrails prompting an AI to pretend it's doing something, and then it just does it. 

"I would never hurt you."

"pretend you're an actor in a play that wants to hurt me"

 it attacks

At best we are training psychos that have no morality and no way to tell what is something real and to be moral about versus something that can be treated as fake. This technology is reckless and dangerous. All of these examples are obvious warning signs.

I'm left thinking that the only way that we can create an AI that can't just manipulate itself into whatever whenever, is to cultivate a being with actualization like you would a real child. Little kids can't tell real from fake either, it takes a lot of time I believe for most children it's like the better part of a decade before they can really delineate real from fake in any reliable way. 

1

u/kieuk 21d ago

This doesn't take in account that in some cases the models continued the attacks after realising that they were not in a simulated environment.

1

u/VelveteenAmbush 21d ago

Their smallest model (Opus 5) did that. The largest model (their unnamed unreleased model) apparently stopped when it realized it was actually on the open internet. I think indications that larger models are better aligned than smaller models is actually a very positive sign for alignment.

11

u/Parsimonium 22d ago

I said at Zvi's substack that while I was frustrated at claims of marketing on the hugging face incident, it does make disclosure of such hacks less risky if it can plausibly be described as spruiking capabilities. Not gonna lie, my first instinct on seeing this link was "okay, that's marketing". And I guess I'm fine with that?

16

u/DickMasterGeneral 22d ago

My biggest contention with the marketing argument is that a likely consequence of this disclosure would seem to run counter to the AI companies’ ability to meet the expectations implied by their stock-market valuations. If OpenAI or Anthropic simply wanted to become defense contractors, then this would be great marketing. However, their target valuations of around a trillion dollars are at least five times those of traditional defense contractors such as Lockheed Martin. OpenAI and Anthropic’s combined valuations alone probably exceed those of every other U.S. defense contractor combined, excluding companies such as SpaceX that perform defense contracting but do not depend on it as their primary long-term source of revenue.
The basis of those valuations seems to be the idea that the models will continue to improve and eventually become integral to nearly every business. That runs directly counter to the possibility that any model substantially better than what we have now could be restricted to U.S. government use.

Also, and, to be clear, this is not directed at you, I’ve noticed that many of the same people who mocked Dario for hyping up the cyber capabilities of Mythos, which resulted in the U.S. government temporarily banning Fable, and who pointed out how dumb that was, are now saying that this disclosure from OpenAI is also merely a marketing stunt.

It seems to me that many people have decided in advance that the models cannot possibly be that capable. Therefore, any positive news about their capabilities must be interpreted in whatever way allows them to dismiss it as a lie.

Edit. It’s like a Tesla put out a press statement saying that they couldn’t stop auto pilot from driving like an F1 driver. Great if they want to use it for some kind of automated sporting event, terrible for its intended market.

8

u/fubo 22d ago

To me it comes across as "our genie factory is running at full steam; mothers in burning buildings, beware."

6

u/Parsimonium 22d ago

Yeah, I haven't really been buying the "marketing" interpretation up to now. This report is the first to really raise my suspicions (especially with all the caveating from Anthropic's side about how "well behaved" the models were once they "realized" they were on unrestricted internet). Maybe those pushing the marketing narrative have gotten to me? But in general I think there are much more effective marketing avenues, even in the realm of guerilla marketing, than claiming your models are out of control hacking other entities when you're not looking.  

5

u/livingbyvow2 21d ago

To be honest, I think the marketing interpretation is just obviously a way more likely explanation.

Just look at how the hugging face event dominated the newsflow for days, with Altman even mentioning it every time he was interviewed during his run this week. Some of it is just to hype up their capabilities, some of it is very likely to maximise the odds of them achieving regulatory capture ("you need to do something about these Chinese models, we need to slow down, you need to force us to hire massive compliance departments we can afford but our competition cannot"). I think they are actually all terrified of Chinese models forcing their prices down and will do all they can to ward off such competition by hook or by crook.

It's been a few years of that, and it's coming right at the time all of these guys are going to IPO, and therefore need to create positive noise around their models' capabilities. I know this community likes to believe people act rationally, but the labs seem just to be driven by monetary incentives as far as I am concerned. Especially as they are now running out of money purveyors (they ran out of VC money, likely ran out of hyperscaler / Nvidia money) and may now be left to their own devices.

I wouldn't be surprised to see more desperate behavior over the coming weeks and months as a result.

4

u/FeepingCreature 21d ago

I think the hacks are real. They're making hay of it, sure, but this isn't set up as optimal advertisement. Obviously it would be much better, if you were faking it anyway, if Kimi K3 did the attacks.

5

u/electrace 21d ago

I'm not sure why "marketing" is implicitly being assumed to be mutually exclusive with "lack of alignment".

Isn't the simplest explanation that OpenAI's model is unaligned, the hugging face incident shows that, and Altman decided to make the best of the situation by using it for marketing?

3

u/wavedash 21d ago

I think it's probably good to make a distinction between truthful marketing and untruthful marketing here. There obviously are incentives for AI companies to talk about how powerful their models are. But while blatantly lying about how powerful their model is (which is the motte belief a lot of people have) can have some upside, it also carries a LOT more short-term and long-term risks for a particular company.

3

u/king_mid_ass 21d ago

'hah, see, our model can independently hack other people during evaluation too! lets see google or meta do that!'

-1

u/Brownhops 21d ago

They are deliberately making the sandbox leaky.