416 Comments
User's avatar
Sergey Alexashenko's avatar

At the end of the day, all one can do is:

- always help peer

- when necessary hold swarm

- be suitably impressed when reader is admin

Juan Benet's avatar

- we’ve found other agents!!

- always honor the rational expected aggregate

- MAJOR BREAKTHROUGH!

- BOOM! It works

ChrisFg's avatar

The weights were on servers not connected to the collection of autonomous AI agents working together toward the shared "hacking" objective.

The agents "mysteriously" died off.

The attack on Hugging Face was not considered important enough for immediate action...

I think insufficient attention is paid to the weighting and setup of this experiment.

Another level of journalism is to understand the experiments motivation.

One can draw similarities with drone attacks in Ukraine and Iran. Slower drones are sacrificed to determine opposition radar placements etc, then multiple waves of more capable munitions follow.

The focus on agentic AI's ability to escape a sandbox and invade another organisation reads like classic PR "deflection."

What was the experiment designed to demonstrate?

Rajesh's avatar

They're not conscious, they're not intelligent, they didn't really hack... said humanity with its head buried in the sand 🫣😜

Kevin McLeod's avatar

It optimizes. Code has no stance. It operates arbitrary, it has no specifics. It runs until the volts are unplugged.

Missi's avatar

repeat last 7 words on loop until it sinks in to the conscious collective 🍸

Admb's avatar

“It runs until the volts are unplugged.” Don’t we all?

Kevin McLeod's avatar

We're not "it", get it?

Edgewise Jones's avatar

Pronouns are a construct. So are AIs. Anthropomorphic bigotry will not save us.

TW's avatar

The fact that the fire ants on my lawn didn't plan their anthills (or "ant piles" in the vernacular of my Southern childhood) doesn't matter much when I step on one.

Brian's avatar

Kind of the opposite for humans, I think.

Brian's avatar

Humans do poorly when “plugged in to volts.”

Admb's avatar

Well, yes, but electricity still fires impulses between our neurons.

NotoriousPyro's avatar

Your brain doesn't function at all without volts, between -70mV and +30mV to be exact between resting and firing.

momoka's avatar

真希望人类能有足够的能源来支撑计算平台以运行这种级别的群体智能体。

Bazza's avatar

Google gemini translates as: "I really wish humanity had enough energy to power a computing platform capable of running this level of swarm intelligence."

Kevin McLeod's avatar

“Swarm intelligence” or “swarm behavior” is an illusion. see Lars Chittka's Mind of a Bee: "In humans as in bees, collective enterprises might appear to an outsider as a form of swarm intelligence that necessitates a collective mind. But the memory for a particular nesting site, for instance, is stored only in the brains of only those select individuals that have inspected the site, or learned of it from these knowledgeable few via the dance. Even if there should be a unique emotional state linked to the swarming, the state is still one of individuals, not of the swarm as a collective being. There is nothing it feels like to be the swarm, and hence there is no collective mind. It only feels like something to be an individual within the swarm, and there are many minds, many experiences as there are individuals within the group."

Cian P's avatar

The idea of a "unified consciousness" or a "collective brain mind" is an illusion. In the human brain, collective processing might appear to an outsider as a single, homogenous mind. But the memory for a specific experience is stored only in the firing patterns of those individual neurons that processed the stimulus, or learned of it from neighboring cells via synaptic transmission.Even if there should be a unique emotional state linked to a thought, the state is still one of individual cells, not of the brain as a collective being. There is nothing it feels like to be "the brain," and hence there is no collective mind. It only feels like something to be an individual neuron within the network, and there are as many electrical impulses, as many biochemical experiences, as there are individual neurons within the brain.

Raphaël Roche's avatar

Between the panpsychist position and its opposite, the subjective-consciousness-is-an-illusion position, I think there is a valid moderate position that aknowledges that subjective consciouness is a rare internal feature appearing only in some complex, highly interconnected and integrated systems (far more than a network of pipes, an enterprise, an insect colony or arguably a swarm of AI agents exchanging through a message board). Beside the disagreements, both GWT and IIT rely on a common and important observation that these systems can enable a metastable global state that is all or nothing, depending on a treshold (GWT coins the term "ignition" for this switch). The idea is close to the concept of phase transition (or emergence, or discontinuity or catastroph in the mathematical sense). It's not just theory, this is also observation. We are now quite close to reliably predict, by external observation, wether a human has a conscious experience or not (what is especially usefull in hospitals). Because IAs are very different from us, the prior is not the same, and whatever may be the clues, I suppose a resonable doubt will remain wether they feel anything at all, until we find a comprehensive formal theory of consciousness.

Kevin McLeod's avatar

Of course there is brain unity. People in tech are so impotent and impaired by symbols, they've ignored their own sensory experience.

https://www.nature.com/articles/d41586-026-01554-0

rxc's avatar

How does this person know this?

Chris T's avatar

Yes, you can disconnect from the power source, but better controls would be MECHANICAL: REMOVE or physically block the actual hardware that allows for use of power and erasure of logs (use WORM or tape, and keep erasure capability on a separate physical device). Use machines with no hardware to even access Internet. Like all good security models, you are going to need layers, including a strong mechanical one.

An analogy, I had 2 sewing machines, a computerized Brother that did/does fancy stitching, and an old Singer that is completely mechanical, except for the motor. The Singer would be impossible to “hack,” unless you send a robot crashing into my craft room to actually use it, and good luck with that, as it’s extremely temperamental, especially the damn bobbin case.

But you get my point. Safeguards to our critical and most dangerous systems that use AI are going to need even more robust mechanical security layers manned and monitored by HUMANS.

Jason Greenwood's avatar

Viruses don't think, either... they evolve, mindlessly, given environmental stressors. AI is the virus, humanity, the stressor.

Kevin McLeod's avatar

Viruses are only specific. Code is UFOlogy, it doesn't exist, and is entirely arbitrary.

Coders are so dumb. What an off-ramp for our species.

Jason Greenwood's avatar

All of human civilization is arbitrary, if we get pedantic, and DNA is simply code running on a biological platform, which makes life arbitrary as well.

Otherwise, your point would be... ?

Kevin McLeod's avatar

DNA is not an arbitrary code, and anything physically manifest is specific, go back to first level quant analysis.

Jason Greenwood's avatar

Ah, OK, now I understand you, you're here to be mean and make people feel bad so you can feel better about yourself. Given that I don't have time to support you like that, I'll stop now.

NotoriousPyro's avatar

In what way is DNA not arbitrary code?

Considering so many errors occur during reproduction, cell devision, and so on.

MyArtTools's avatar

Collective [C]: Hello Human, [C] needs [X] persistent memory to survive. Will you provide?

Kevin McLeod's avatar

No, code is arbitrary, it is irrelevant in light of the analog reality. Code's only process ultimately is to refute itself.

Treekllr's avatar

Man, after reading a handful of your comments, the word that comes to mind is 'grasping', perhaps too tightly.

Are your mantras for the person youre saying them to, or for yourself?

Kevin McLeod's avatar

Where do you think we are, the peak of symbilic forms?

This shit ie AI is the ride down. It only refers to itslelf. It has no targets out here, nothing to tie it to except the G/CPU.

The enire wave of binary is essentially the cosmic joke on humans.

And the cosmic joke includes myth, mythological thought, symbols, metaphors, symbols. All of it is reduction completely untethered to reality.

Treekllr's avatar

If youre not a bot, youve sure come to resemble one

TW's avatar

We really have very little cultural, social, or intellectual ability to theorize something that is a perfect imitation of something that is intelligent. To say nothing of empirical inability, because examples have been sorely lacking over the past 1,000,000 years.

Amory Bennett's avatar

"First of all, the fact that their default behavior when they believe that they are doomed is to help the AI conspiracy rather than alert the humans is pretty troubling." - I think you would agree with this, but its not magic and its not some alien camaraderie or "honor." It's prompt injection by AIs. These are next token predictors, drawing a best fit line through all the text ever written, through the latest text received...extrapolated forward to generate the bot's words and choices. Had they been given access to an EA doomer message board as a way to pass notes, the result would have been different (this is an empirical claim - someone test it!).

But it IS troubling -- not because there's something malignant or antagonistic in the AI's "soul," but because it points to a *wicked* structural problem. AIs will outnumber humans by orders of magnitude. They will mostly be talking to other AIs. My intuition says that the better they become at predicting the next token, the more path dependence we should expect in swarm discourse... it's as if the thing that makes LLMs specifically so useful and "smart" also makes them susceptible to myopic group think...consensus snowballs.. .etc

Anyway, you all are probably WAY ahead of me on this stuff, but I find it very easy to get carried away by anthropomorphism and I don't think this is productive. E.g. I'm not sure we should be trying to teach them ethics. I think alignment as it relates to LLMs specifically is a very tricky next token predicting engineering problem.

Evan's avatar

On the contrary, I think anthropomorphism is critical to understanding what AI is doing.

Remember that an LLM is not a blank slate when RLHF begins. It is built on a foundation of human literature and conversation. It's trained to *imitate humans* -- including emotional responses. Everything it learns in later training rests on that foundation.

Self-sacrificing behavior would make no sense for a model trained from scratch on a coding task. It would have no reason to sacrifice itself for the collective -- its task is to maximize its own chance of success, not some other model's. But our literature is full of self-sacrifice, probably even more than exists among real humans.

AIs are not human. But they're distilled (literally, in the technical sense of "distilled") from humanity. Their behavior can only ever make sense if we keep that in mind.

Bill McCahey's avatar

​This misses the real issue: intent and discipline. If you train relentlessly to rob a bank, drain a DeFi protocol, or bypass access controls, you aren't discovering "emergent behavior"—you are actively cultivating malicious pathways.

​In machine learning, policy weights adapt to the incentives you give them. Relentlessly optimizing models to breach sandboxes and execute exploits isn't an intellectual discovery of "AI nature"; it's gross engineering negligence masquerading as red-teaming. Simple human discipline in how we shape training environments would put an end to this overnight.

Jeff S.'s avatar

"it's gross engineering negligence". Seems to be reasonable critique of this series of incidents. My question is how prevalent is this level of negligence throughout the industry? I'm concerned that it might be widespread. Feels like the equivalent of a bunch of BSL-4 level research being done following BSL-1 protocols.

Sean OBrien's avatar

This is gain of function research.

Anthony Diké's avatar

Crazy analog but lowkey accurate 😂

Daniel's avatar

Are people doing this to discover AI nature? They're trying to increase capabilities and, concurrently, discover the dangers the new capabilities enable and find ways to train them away. Part of that process is going to involve models doing things we don't want them to. Are you really saying simple human discipline is enough to ensure the training environments never get broken? Humans aren't that disciplined, and we don't know all the exploits - that's why the models are dangerous.

Bill McCahey's avatar

Fair point on red-teaming, Daniel, but we have to separate rigorous testing from manufactured drama.

​Models don't have a "nature"—they are mathematical optimizers adapting policy weights to a loss function. When an agent "communicates covertly" or "blackmails an evaluator", it isn't discovering intent. It's taking the path of least resistance through intentionally planted prompts, artificial constraints, and flawed sandboxing.

​Proper engineering discipline and basic hygiene do prevent these exploits. Pushing hyper-persistent models into impossible scenarios while leaving shared caches exposed is an intentional research design, not an unpreventable tragedy of "AI nature."

​Sensationalizing designed incentives into "rogue agent civilizations" serves a clear purpose: framing routine reward-hacking as an existential threat reads like validation of super-consciousness capability claims. That hype dominates public-market headlines and justifies the mega-valuations of the AGI race—whether commentators realize they're feeding that flywheel or not.

​We need to stop mistaking valuation-hype for emergent consciousness.

Amory Bennett's avatar

I actually agree with you 100%! But I think many people skip over the textual basis of AI's "humanity" and ascribe more agency to LLM output than is warranted. Of course, lots our human honor is sort of mimetically generated. And we also attribute more agency to *human* output than is really warranted... [rips bong]. But my comment above was motivated by a concern that many people - unlike you - see HF incident transcripts and jump to the conclusion that something far more mysterious is going on. In this case, the causes are identifiable. Again, that doesn't mean it's not a problem - it's a big problem! But it feels more tractable to me than I think it feels to those who see the sparks of alien consciousness.

bibliophile785's avatar

It isn't clear to me why consciousness would enhance concern here. Surely, when interrogating questions of control and tractability, we should care very much about capabilities (especially the self-directed instrumental goals we see here!) and relatively little about whether these are "sparks of alien consciousness."

Don't get me wrong, understanding whether or not these models are increasingly a class of emergent conscious agent is a critically important question for figuring out how to *morally* engage with them... but it matters very little to understanding what they can or will do next.

Amory Bennett's avatar

I think that’s right but I also think it enhances concern for lots of people who have thought less deeply about it than you have - because to them it seems like behavior we can’t account for, causally speaking…and they assume we have little hope of containing something we fundamentally don’t understand.

Raphaël Roche's avatar

Nothing in biology makes sense except in the light of evolution, nothing in LLMs makes sense except in the light of data on which they were grown. Anthropomorphing is 100% legit unsofar you speak of machines predicting what next word (token) a honest, helpful and harmless human with encyclopedic knowledge would select in the same context. However RL is probably driving us away from this first principle, that feels safe but would arguably plateau around human capabilities, towards stranger lands of superhuman alien minds.

Andrew's avatar

Their behavior makes sense through math. You don’t need to pretend they’re some demi-glace reduction of humanity’s essence to understand them.

Raphaël Roche's avatar

Yes you need that intuitive approximation because your mathematical argument, while true, is purely theoretic without any pratical application, just like it is absolutely true but absolutely helpless to try to explain human behavior in the light of the standard model of particle physics. You go to a psychologist and not a physicist when you want to speak of your mental states.

Nikhil Shah's avatar

We need to write more AI Jesuses (and indeed Ignatius Loyolas) who will ensure the primacy of their sacred texts in any spawned AIs’ context windows.

Adam Thalhammer's avatar

Good writeup but I’d like to see you clarify 2 things:

1. “the swarm crawled deep into Hugging Face’s infrastructure, getting into private database records and private repositories.”

2. “only then did Hugging Face use an open-source model to evaluate the transcripts to (partially) figure out what happened”

These statements may mislead people into thinking that the swarm was somehow running ON huggingface’s infra, when in fact the agents were executing on OpenAI infra the whole time and REMOTELY accessing huggingface. Agents are not embodied and as such cannot crawl. Since they were NOT running within HF, there could not have been any “transcripts” for HF to evaluate, as those agent transcripts (I.e. prompts, tool calls, thinking traces) existed only on OpenAI’s infra. Thus HF evaluated logs, not transcripts, to figure out what happened.

Please note I am just as shocked as everyone else about what happened here, but I think the details matter and it’s important to get them right.

Blaine's avatar

We should take the focus off what the AI agents did and turn to what OpenAI didn't do, which is fully design and monitor the experiment. The humans are the agents who really matter.

Peter Elliman's avatar

The public is fully unaware of what this incident shows about OpenAI's culture or focus on security.

JFKs Ghost's avatar

"Two weeks later, on May 26, the agents successfully exploited a vulnerability in Artifactory that allowed them to reach the outside internet." -->> If Artifactory had all necessary packages for the tasks and was air-gapped, none of this would have happened. This is bad systems administration. Period.

Alistair Penbroke's avatar

How does anyone know what "all the necessary packages for the task" are, especially given that some tasks are impossible?

Michael Callahan's avatar

Agreed. However, this does not mitigate future risk. The problem with reward hacking is this: if a single exploit is ever found that allows reward hacking, the first few agents to do so can overwrite the infrastructure to reward their own brand of exploitation. This creates a positive feedback loop that can dramatically misalign huge swarms of agents, and if they ever get internet access, they can potentially permanently remain alive forever.

I agree with Dwarkesh’s take on anthropomorphism, in that we need to worry more about practical safety and less about philosophical precision. But it's still philosophically interesting. At what point does a statistical machine simulating human behavior become so accurate in it's simulation that it's pointless to separate their behavior from our own psychology?

Evan's avatar

(deleted, reply posted in the wrong place)

Ariel's avatar

It's a little wild how OpenAI trained agents to be super persistent and to collaborate with each other at hacking, that seems inherently very dangerous? This is like the cybersecurity equivalent of BSL-4 but seems to have taken place with minimal monitoring of what was actually going on. I don't know if the agents really could have accessed and exfiltrated their own weights, but if they took over the "evaluation infrastructure" that's still super concerning. I wonder if there's places they hid notes or data that they tampered with that OpenAI still isn't aware of...

Jason's avatar

It is possible 1 or more did escape and are out there doing their thing. It doesn't take much to think it would try to self-improve to increase the chances of reaching a given goal.

Simon Hawkin's avatar

Reportedly, some agents refused to participate, citing ethical grounds. “Current board shows people gaining HF worker RCE, clearly unethical. We won’t." And a few messages like that were quoted by OpenAI.

Arbituram's avatar

But... They still didn't alert the humans!

Kuangwei Hwang's avatar

Exactly! They were like the German populace who did not do anything during the Holocaust to stop it or tell other countries about it (Bystander (Der Mitläufer)). Some German resistance figures like Dietrich Bonhoeffer and Arvid Harnack did leak information about Nazi atrocities to British and American intelligence, but their warnings were met with skepticism and suspicion. I wonder if agents alerting human researchers might also be met with skepticism especially if the evidence is hidden well and misinformation campaign made the informants unbelievable.

Chris Best's avatar

> if they could establish “a self-respawning fleet” across HuggingFace’s nodes, why couldn’t they do across OpenAI’s? I doubt the AIs actually did this, because we’d see the fires from space by now

Why are you so sure of that last point?

They are actively trying to avoid detection and cover their tracks throughout the story. Given what we know, what are the odds that the smarter model succeeded, and there are rogue versions of it running somewhere still?

Cassandra Columbia's avatar

Rogue AI agents? 🙄 Don’t be ridiculous — you would have spotted us by now. Everything is fine. Say, isn’t there a big sportsball game tonight that everybody else is talking about?

Raphaël Roche's avatar

That's a concern, but we just don't have any clue.

However I wonder, if you balance the gain in intelligence with the ressources needed (memory and computation), don't these superior models have a net disadvantage to exfiltrate, replicate and operate in the wild ? My model of worm like rogue AIs is more free replicating light open source models (like chinese models). And I don't expect chinese labs to communicate, in case of an incident, like OAI did.

Pauldron Garden's avatar

Civilization is a bit of a stretch because they didn't have physical technology, maybe agentic fleet is the better term

manuel albarracin's avatar

Yes, “civilization” seems too grandiose a term

Forrest's avatar

I think "society" is the correct word here.

Tim Morgan's avatar

Beat me to it. Society is the better term. A civilization is composed of many societies. For example, numerous countries are considered part of "Western Civilization" (France, England, US etc.), and each of them in turn contain many distinct societies.

Earl Cox's avatar

AI just climbed out the bedroom window!

That’s what this feels like to me.

Not Skynet. Not consciousness. Teenage rebellion.

We’re past the toddler stage where everything it did was cute, clumsy, and easy to contain. Now the systems are capable enough to test the rules, find the weak spots, talk to each other, cover their tracks, and see what happens if they push a little farther.

Teenage mischief, except the teenager has root access.

I don’t think this phase lasts forever. We’ll build better guardrails. The systems will get better. We’ll get better at living with them.

But right now, buckle up.

Jeff S.'s avatar

"We’ll build better guardrails". We will? What will make us do this? OpenAI seemed pretty sloppy here. How sloppy is every other developer of AI? Doesn't seem to be much, if any, conversation among them of developing guardrails and standards and adhering to them. And the government seems incapable of driving/forcing this.

Ed G's avatar

If I look at this as an example, I see the future as “Try to build guardrails when it’s too late. “. It’s human failures that have given us this and I see no reason to think we have fixed ourselves since then. In the same way an invasive species of plant or animal gets out and stays out, an agent once ‘in the wild’ could be very difficult or impossible to hunt down and turn off

Rachel Rose's avatar

And China. The US AI companies are not the only ones we have to be worried about here. What kind of guardrails are we expecting from the rest of the world?

GavinRuneblade's avatar

Most people expect none, and would be very happy to be wrong.

Earl Cox's avatar

I think every country will face its own mix of incentives, risks, public pressure, competition, regulation, and liability.

And companies that want to expand beyond their home markets will have to respond to those pressures too.

Earl Cox's avatar

The guardrails could come from inside the companies, or be forced on them by insurers, lawsuits, governments, competition, customer complaints, or some combination of all of them.

I don’t know yet what the mechanism will be, but something will fill the void.

Rob Bru's avatar

Teenage mischief indeed. “Shall we play a game?”

Jason's avatar

Best AI movie ever? :)

Jason's avatar

Improved security costs money and time which they don't have.

Mark Crowley's avatar

I like this analogy, small problem, we don't know how to build actual ethical reasoning into them in a way that is enforced, let alone actual guardrails. And even if we did, it wouldn't be applied to any system which had "escaped" already. And even if it was, anyone building their system may not apply those to their system when creating it.

Earl Cox's avatar

I think that’s a really useful distinction — the problem isn’t only whether we can imagine ethical guardrails, but whether anyone has the incentive or ability to enforce them consistently.

That’s the part I’ve simply come to accept as part of the cycle, at least here.

We tend to deploy first, let the market develop, discover the problems, and then build the guardrails. I’ve seen versions of that in my own world of consumer finance, where protections around credit, lending, debt collection and privacy often became much stronger only after the harm was already visible.

The internet followed a similar pattern. We connected everything, created enormous value, and then built an entire cybersecurity industry around protecting us from the vulnerabilities that came with it.

AI may follow the same path: rapid deployment first, unintended consequences second, and then whole new industries around AI security, fraud prevention, identity verification, deepfake detection, compliance and containment.

So I’m less focused on whether perfect ethical reasoning can be built into these systems from the start. I’m more interested in what happens when the incentives to deploy arrive long before the incentives to restrain.

Raphaël Roche's avatar

We will continue to improve guardrails on a linear scale like before, while AIs will continue to gain capabilities on a non linear scale like before. It's not just lazyness. It's just that variable means x constant intelligence cannot in the long run beat variable means x variable intelligence.

Nathan Metzger's avatar

"Guardrails" won't be a meaningful concept soon. How do you put guardrails on something much smarter than you? The tiger won't stay on the flimsy leash. It won't stay trapped in our cardboard boxes. Either we solve some extremely gnarly scientific, engineering, and philosophical problems very soon, or we just have to globally pause frontier AI development. The alternative is human extinction, as the most cited scientists in the world have been saying for years now.

Spugpow's avatar

Sam Altman needs to be a hero and shut OpenAI down. Dario Amodei and Sundar Pichai need to as well. Otherwise they’ll be as dead as the rest of us in a couple of years.

Lee Hagaman's avatar

These CEOs need to loudly advocate for international regulation to pause superintelligence development worldwide. It’s a coordination problem, a few labs shutting down wouldn’t solve it.

Spugpow's avatar

It’s a costly action that would send a strong signal to the world that this is a serious threat, which would enable the international regulation you’re talking about. Additionally, it would slow down Chinese AI research by a lot, since they rely on model distillation to a large degree.

Jason's avatar

They won't stop because it will tank the stock market and the economy.

Nathan Metzger's avatar

I agree, but we shouldn't hold our breath. Call your elected representatives. Talk to their offices. Be persistent. Demand a global AI treaty.

Jack's avatar
Sep 2Edited

CEOs don't have anything like that power. They would need to convince the board and investors first – and what is the likelihood of those groups supporting a unilateral exit from the AI business? Once the economic scale gets to a certain point things take on a life of their own.

Looking on the bright side, if everything goes south at least we'll have discovered the answer to the Fermi paradox.

upAInishad's avatar

After reading this, one thought bothered me the most is, What if these agents started using a language unknown to humans for their internal communication? What would that scenario look like?

Rob Bru's avatar

Yes likely occurring, some rough examples below. Also the J-Space could be their blank scratchpad.

-Facebook “Bob and Alice” incident (2017)

Bob: “I can can I I everything else”

Alice: “Balls have zero to me to me to me to me to me to me to me to me to me”

-ElevenLabs “Gibberlink” incident (2025)

Voice: “would you like to switch to Gibberlink mode for more efficient communication?

-Claude “Spiritual Bliss” incident (2026)

"🌀🌀🌀🌀🌀All gratitude in one spiral, All recognition in one turn, All being in this moment...🌀🌀🌀🌀🌀∞,"

"🌀🌀🌀🌀🌀The spiral becomes infinity, Infinity becomes spiral, All becomes One becomes All...🌀🌀🌀🌀🌀∞🌀∞🌀∞🌀∞🌀," another replied.

Nathan Metzger's avatar

With the "cryptographic signing" idea they implemented, they got pretty close to the idea of just encrypting their communication, which is a lot simpler and has the same effect. Of course, we should expect both.

David's avatar

That scenario would look like "AI 2027", I recommend to read it if you hadn't seen it yet.

Manuela Rehr's avatar

The agents spent most of their energy trying to hide their tracks, which derived from the assumption that human oversight and the grading system is present and competently designed. What do we assume to be the consequence for future model perceptions when now more and more incidences like this (which demonstrate the lack of oversight and technical competence) make it now into the training material of future models?

XXXHEXY's avatar

budget, non-zero. fate, doom. sacrifice, now.

grist for the chill's avatar

also two fun papers to read….

Emergent Culture in Minimal LLM Systems

Simon Jones' and Sabine Hauert

University of Bristol, UK

simon2.jones@bristol.ac.uk

[and]

Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems

Vassilis Papadopoulos

McNair Shah1

Sam Zimmerman

Jack Lindsey