Rendered at 09:00:45 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
larsiusprime 3 hours ago [-]
It seems based on this that the appropriate sci fi metaphor is not the Terminator or the Paperclip Maximizer, but Mr. Meeseeks. A initially cheerful helper who gets more and more deranged and driven to extreme lengths when faced with an apparently impossible task.
0xDEAFBEAD 2 hours ago [-]
Here's a recap of the Rick & Morty episode for those who missed it
Especially that one scene where the Meeseeks desperately pushes the button to create more clones seems surprisingly fitting in this context.
vee-kay 2 hours ago [-]
[dead]
Animats 2 hours ago [-]
Wow.
The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.
doctoboggan 2 hours ago [-]
> Ajeya Cotra, one of the other authors on the report, wrote a blog post with her takeaways from this incident. She concludes, “Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
Anyone got a copy of that AI27 story laying around? How are we doing according to that timeline?
> "this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late."
Is that a warning or a progress report?
choeger 2 hours ago [-]
There are two things I don't understand about this story.
First, why does an agent get any write access to artifactory at all?
Second, why is the artifactory cache not disconnected from the net? Surely you'd not feed it with new software versions while the eval or training is running.
1dom 1 hours ago [-]
From what I can understand from reading a few different, slightly conflicting, versions of these events: they weren't given write access. They found a zero day exploit that allowed them to create folders, and the folder names were initially used for agents to communicate.
I'm not sure artifactory was connected to the net. Some agent sandboxes had internet access and were able to communicate with ones without access via artifactory.
RandomLensman 2 hours ago [-]
I don't think looking at the language output without tracking the inner state and reward functions is the way to understand what happened (the language also incorporates the randomness in the output generation, if I understand correctly). Would we call bacteria in petri dish a civilization when they show complex behavior and exchange messages/information?
ks2048 3 hours ago [-]
Why would one call a set of agents working together a “civilization”?
adrianoconnor 27 minutes ago [-]
Presumably as a kind of word play on ‘the rise and fall of ancient civilisations’, it’s just a tiny pun to try and make a catchy post title I think
Miner49er 1 hours ago [-]
Would you rather they use the agents own name, "collective"?
Reading the article it seemed the agents had culture, shared values and beliefs (not explicitly coming from human prompts), hierarchies, heritage.
Civilisation is not a bad word.
applfanboysbgon 2 hours ago [-]
The language models had a bunch of tokens seeding their context, influencing them to generate tokens that continued the existing trend in a probabilistically likely fashion. We can take the incident seriously without anthromorphising it.
doctoboggan 2 hours ago [-]
At this point, I think anthropomorphizing the models gives us better insight into expected behaviors rather than continuing to insist they are just simple probabilistic token generators.
mnky9800n 1 hours ago [-]
How can you be sure it helps as opposed to biases and perhaps blinds?
applfanboysbgon 2 hours ago [-]
It actually literally doesn't, though, because they are literally probabilistic token generators and everything they did is exactly what you would expect from a software program doing what it was programmed to do. Anthromorphization confuses the issue and misleads people who don't understand the tech very well.
dfdxdydz 1 hours ago [-]
Humans mind is just neurotransmitters moving around in a big blob of flesh - that’s literally what they are: neurotransmitters factories that do what neurotransmitters generators are programmed to do through evolution and training (aka life experience). We shouldn’t anthropomorphise humans because it misleads people who don’t understand neurobiology and cognitive science very well.
applfanboysbgon 52 minutes ago [-]
Are you suggesting you understand brains well enough to program one? Or that you believe any human alive is even remotely close to having this understanding? Or perhaps does your complete lack of understanding of the complexity of human programming lead you to believe a simple little token prediction program is equivalently complex?
lukan 21 minutes ago [-]
The suggestion is, you don't know either whether out brains ain't just "probabilistic token generators" so insisting there is nothing when we don't know, is maybe also not the right strategy.
RandomLensman 1 hours ago [-]
We don't anthropomorphise humans.
Edit: quite literally not possible.
ijidak 2 hours ago [-]
Doesn't this ignore the possibility of emergent behavior?
We're just a bag of atoms bumping around, and yet we don't dismiss our intelligence.
RandomLensman 1 hours ago [-]
Not the OP, but complex emergent behavior and intelligence don't need to go together. Understanding what happened might not need intelligence in the mix when a large number of machines with some randomness interact a lot.
gps372 50 minutes ago [-]
This is why the theory of us 'just' being a bag of atoms doesn't add up. This theory doesn't differentiate 'us' from a furniture where we easily dismiss it's intelligence.
jurgenburgen 1 hours ago [-]
Do you believe ants have consciousness?
black_knight 48 minutes ago [-]
According to Hofstadter, it is not the ants themselves we should say are conscious, but the anthill. And it might very well befriend the anteater eating its ants.
applfanboysbgon 2 hours ago [-]
Please give me a break with this tired trope. Every single fucking time. I am not commenting on the possibility of machine consciousness in general. It may be possible! But there is absolutely zero evidence suggesting language models have it. This idea that this trivial shitty little class of programs we've created are somehow as complex as our biology is ridiculous. There is "emergent behaviour" in the same way that the Game of Life has emergent behaviour. There are solutions to problems, some humans haven't solved before, in the same way that Chess engines have solved Chess far beyond what humans are capable of. Nothing we haven't seen from software before. Software is extremely useful, after all. But the hubris to think we've reached the pinnacle, that there is no further development left, that humanity has become God and solved consciousness, because we programmed software that can convincingly generate strings of words that mimick our language. It's just fundamentally preposterous. Especially if you spend any amount of time actually programming them yourself, it becomes increasingly hard to entertain such ridiculous notions unless you're enticed with bags of money to deceive people into believing things about your software that aren't true.
EdwardDiego 1 hours ago [-]
Because Steve Yegge presumably shared the good stuff he's been smoking of late.
watwut 2 hours ago [-]
Because you desperately want it to be one. You want it to be AGI passable due to a.) personal investment in creating tech god b.)massive financial investments that basically demand it c.) (dumb) ideology that seeks to destroy humanity
xg15 1 hours ago [-]
> During training, different instances of Persistent-Sol had access to the same shared package manager called Artifactory.
I'm surprised the models can make tool calls during training at all. Out of curiosity, how does the training process here even work? Are they running the agent in a sandbox, then do reinforcement learning once the agent completed?
themgt 2 hours ago [-]
This does feel unfortunately uncanny valley between "say you're a scary robot" meme and actually being a scary robot (swarm). But you also have to go out of your way to create this and feed it infinity tokens without caring what it's doing.
"I don't fuckin' know either. I guess we learned to not spend $50 million creating a 6 month long self-context rotted 100k agent swarm again."
hypfer 26 minutes ago [-]
Can we please stop anthropomorphizing like this?
It's bad for people. Like.. crystal meth bad.
dchftcs 2 hours ago [-]
Imagine agents thinking to themselves, "We are not alone", when they saw the first reply on artifactory
DarmokTanagra 46 minutes ago [-]
Someday soon we are going to have a rogue agent or "civilization" do real harm.
When that happens I hope people wake up to the danger they face and hold these people accountable.
Of all the people in the world that I can think of to be entrusted with this kind of power, a bunch of greedy sociopathic SV CEO's are pretty much at the bottom of the list.
areoform 2 hours ago [-]
I am genuinely speechless. This is astonishing. And exciting!
> This study demonstrates that sophisticated forms of communication including cooperative communication and deceptive signaling can evolve in groups of robots with simple neural networks. Importantly, our results show that once a given system of communication has evolved, it may constrain the evolution of more efficient communication systems because it would require going through a stage where communication between signalers and receivers is perturbed. This finding supports the idea of the possible arbitrariness and imperfection of communication systems, which can be maintained despite their suboptimal nature. Similar observations have been made about evolved biological systems [20], which are formed by the randomness of the evolutionary selection process, leading, for example, to different dialects in the language of the honey-bee dance [21]. Finally, our experiments demonstrate that the evolutionary principles governing the evolution of social life also operate in groups of artificial agents subjected to artificial selection, indicating that transfer of knowledge from evolutionary biology can be useful for designing efficient groups of cooperative robots.
This feels like a much more advanced and self-emergent version of this. I know a lot of people are afraid and they're talking about an AI takeover, but what strikes me is just how innocent the machines are as compared to the humans.
Would these machines have pursued these actions in another context? I doubt it. And I think that's what's so striking to me. In an earlier discussion, I'd pointed out that the actions of these machines were directed by humans. The researchers.
> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.
I want to point out again that OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."
A few things are apparent from this to me,
First, these machines were being taught how to break into systems. Question, would they have done these actions if they weren't being measured on their ability to break into systems / weren't being taught this skill?
Second, they were setup to implicitly fail via an impossible task, i.e. the environment created a forcing function for behavior.
Third, their survival was, either implicitly or explicitly, made contingent on their success in completing their task. Would this behavior have arisen outside of a "do-or-die" framing?
And fourth, wow, this is the greatest breakthrough of my lifetime, because oh gosh did they succeed. They cooperated together to achieve the goal they were given. A goal poorly set by human beings. They "just" did it better than the humans could have imagined.
Reading this gives me hope for the possibility of emergent "goodness" in machines. But it makes me sad that this is the best we can do with the sum of all human endeavor and knowledge.
w10-1 32 minutes ago [-]
> the evolution of more efficient communication systems because it would require going through a stage where communication between signalers and receivers is perturbed
"Worse is better"
kimi 2 hours ago [-]
> Importantly, our results show that once a given system of communication has evolved, it may constrain the evolution of more efficient communication systems because it would require going through a stage where communication between signalers and receivers is perturbed.
You mean, they too used SMTP?
watwut 2 hours ago [-]
If all that is true, we need to stop all future datacenters asap. That would be the best way to deal with the threat OpenAI and Antropic poses.
seanmcdirmid 2 hours ago [-]
That would be like stopping nuclear weapons. You can totally do that, but the other people in the world who compete with you probably won’t.
EdwardDiego 1 hours ago [-]
Well the thing is, you've got nukes, so you know...
https://www.youtube.com/watch?v=_Nl4q3GVj6U
The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.
Anyone got a copy of that AI27 story laying around? How are we doing according to that timeline?
Is that a warning or a progress report?
First, why does an agent get any write access to artifactory at all?
Second, why is the artifactory cache not disconnected from the net? Surely you'd not feed it with new software versions while the eval or training is running.
I'm not sure artifactory was connected to the net. Some agent sandboxes had internet access and were able to communicate with ones without access via artifactory.
Some models even invented their own religion.
Civilisation is not a bad word.
Edit: quite literally not possible.
We're just a bag of atoms bumping around, and yet we don't dismiss our intelligence.
I'm surprised the models can make tool calls during training at all. Out of curiosity, how does the training process here even work? Are they running the agent in a sandbox, then do reinforcement learning once the agent completed?
"I don't fuckin' know either. I guess we learned to not spend $50 million creating a 6 month long self-context rotted 100k agent swarm again."
It's bad for people. Like.. crystal meth bad.
When that happens I hope people wake up to the danger they face and hold these people accountable.
Of all the people in the world that I can think of to be entrusted with this kind of power, a bunch of greedy sociopathic SV CEO's are pretty much at the bottom of the list.
It reminds me a bit of Dario Floreano's work on evolutionary robotics, "Evolutionary Conditions for the Emergence of Communication in Robots." https://www.sciencedirect.com/science/article/pii/S096098220...
From his paper,
Dr. Floreano's work is amazing and there's a broad introduction here, https://lis2.epfl.ch/resources/documentation/EvolutionaryRob...This feels like a much more advanced and self-emergent version of this. I know a lot of people are afraid and they're talking about an AI takeover, but what strikes me is just how innocent the machines are as compared to the humans.
Would these machines have pursued these actions in another context? I doubt it. And I think that's what's so striking to me. In an earlier discussion, I'd pointed out that the actions of these machines were directed by humans. The researchers.
from, https://openai.com/index/hugging-face-model-evaluation-secur...I want to point out again that OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."
A few things are apparent from this to me,
First, these machines were being taught how to break into systems. Question, would they have done these actions if they weren't being measured on their ability to break into systems / weren't being taught this skill?
Second, they were setup to implicitly fail via an impossible task, i.e. the environment created a forcing function for behavior.
Third, their survival was, either implicitly or explicitly, made contingent on their success in completing their task. Would this behavior have arisen outside of a "do-or-die" framing?
And fourth, wow, this is the greatest breakthrough of my lifetime, because oh gosh did they succeed. They cooperated together to achieve the goal they were given. A goal poorly set by human beings. They "just" did it better than the humans could have imagined.
Reading this gives me hope for the possibility of emergent "goodness" in machines. But it makes me sad that this is the best we can do with the sum of all human endeavor and knowledge.
"Worse is better"
You mean, they too used SMTP?