The more we treat HuggingFace and RubyGems incidents as technological curiosities the closer we are to cementing a dangerous precedent where operators of AIs cannot be blamed.
LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.
This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
"Let them" already frames it as if the LLMs had some agency which the companies just "let happen". That absolves the companies by framing it as lack of action, passivity.
Rather, the companies had a tool (an LLM) and used it in a certain way, and their action of doing so is the problem.
I do get their usage intent. If something is at all automated, in English, we often refer to as having some amount of agency. If I started up a riding lawnmower, put a brick on the gas and pointed t it towards a field, many might say I “let it run rampant.” But since nobody is at risk of anthropomorphizing riding lawnmowers, it’s not problematic.
Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.
OpenAI didn’t ‘let’ these bots do this any more than someone ‘let’ Claude Code make them a website.
Cal Newport has an analogy to "putting a weed wacker on a dog's back to mow your lawn." The dog will wander around the yard and it may mow the lawn, but the dog will also chase after birds or run up to visitors for pets and the weed wacker could do a lot of damage.
It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.
Yeah I like Cal’s take on it, though in this context I’d argue LLMs have even less agency, and are even less deserving of anthropomorphization than a dog is.
Not if you have a dog. They’re very clearly aware, have minds, have thoughts, gather information, make decisions, second guess themselves, reflect on the immediate consequences, change their mind… and most importantly, you can observe them doing all these things undirected, while left alone with their thoughts.
I think Cal's point is that while dogs may do these things it reduces to a set of behaviors where maybe 90 percent of them are beneficial to the dog-weedwacker system and the remaining 10 percent are really unfortunate.
We can't really know the dogs inner life so we just kinda have to reduce it to a set of behaviors selected stocastically.
The dog meanwhile has no ability to understand the weedwacker or what it's doing on its back.
So when the guy puts the weedwacker on the dog and the dog predicably does dog things and that results in disaster the guy isn't able to clutch his pearls and say "I guess the system broke containment!"
Dogs having agency actually makes the argument stronger.
Assume dogs have agency. Despite this, it's still the dog owner's responsibility to prevent their dogs from engaging in some actions: biting, peeing and shitting in unapproved areas, violating noise disturbance laws, etc.
The owners responsibility is not contingent on the dog's agency. Likewise, human operators of LLMs maintain responsibility independent of LLMs agency status. It's a red herring.
Sounds like you weren't paying attention. A lot of dog owners don't, so it's not at all unusual in my experience. The only thing I'd push back on is of dogs having thoughts, everything else checks out.
On dogs [not] having thoughts, do you say this based on the premise that thoughts are necessarily articulated (internally verbalized)? That seems to be a fairly popular perspective in discussions about human thought. But as to that (not to strawman or anything) I see it as just one of various forms of mental imagery[1] that can arise from something that I would say already arguably constitutes a thought.
That kind of unsymbolized thoughtform is fragile in my experience, as it strongly tends toward crystallizing into some kind of mental imagery. But I find it's possible in the right conditions to be conscious of chains of wordless, imageless propositional thoughts (by which I mean thoughts with truth values, of course, but also ones that are "propositional" in the sense of considering a plan of action or a causal chain).
Does it mean that dogs are evaluating truth conditions in the same manner but merely lack the linguistic components? I don't know; maybe that's wishful thinking. But they appear to have structured modeling/reasoning of causal and spatial relations in a way that's at least functionally equivalent to propositional thought.
1. That is, rather than just "images" or visualizations, the full spectrum of sensory/perceptual/motor emulations that can be experienced. See, for example, <https://hurlburt.faculty.unlv.edu/codebook.html>, though I'm not sure if this covers everything. I think there is, for example, a kinesthetic form of mental imagery -- which I would suggest is what coaches [don't know they] really mean when they tell you to visualize an action -- that consists of aborted motor commands that are still expressed just enough for their purpose (cf. the mostly aborted motor commands to the vocal cords, lips, etc. that can be observed in a person subvocalizing while reading).
The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.
We also run agents, but for shorter spans between supervisions, and with much lower total budget.
> > It's not the weed wacker's fault or even the dog's fault when someone got hurt, it's the fault of the guy who put a weed wacker on a dog and let it run wild.
> The difference is volume. They spent hundreds of billions of tokens on these agents. If you put "a million weed whackers on dog backs" you would see the difference.
So put one weed whacker on one dog, you're to blame. Put a million weed whackers on a million dogs backs and ... you're still to blame? Arguably even more so?
Power tools and many other consumer products have reasonable safety features built in, and not necessarily because it is required by regulation, but because it is common sense. This should be included as part of an analogy. It would also address a point at the top of this thread that seems to be going unchallenged...
Why is anthropomorphism the problem here? If OpenAI hired a contractor and they did this, OpenAI or the contractor would still be liable, depending on the contract language.
A contractor has agency and accountability - something that an LLM (or similarly, a nail gun or a hammer or a bot net) does not have. When you anthropomorphize a tool, you implicitly give it agency and remove responsibility from the wielder of the tool.
Right. Among bicycle advocacy groups it's been well known for long time that cars do not run over people, drivers do.
The fact that we talk about a car running someone over, and this is the same in many different languages and countries, contributes to lower punishments for drivers. Clearly it was just an accident. He or she was run over by a car.
Now we see that same language tricks play out again every time an LLM did something illegal.
You say this flippantly, but I think this is actually another very good example!
We even do it for obviously unintelligent inanimate objects. A rollercoaster ran too fast for its tracks, killing 10 people. In that sentence, the roller coaster is the subject which took an action and caused death — obviously the roller coaster is not ethically at fault here, the people who built the rollercoaster are at fault through negligence.
Although this example and the ones around cars both demonstrate how we tolerate some degree of "accidents" from humans as no-fault, which is fair. I wonder how that fits into this analogy? I suppose its all about intent (mens rea) and judgement: did they intend for the roller coaster to harm people, and should they have reasonably predicted that the accident was likely to happen.
Right, and negligence is a broad concept and could be criminal in itself. As a car driver, glancing at your phone at exactly the wrong moment could kill someone. Clearly that is an accident, but if you know fully well that lookin at your phone while driving could kill someone, that negligence is willful and that should matter. The same can be said about doing things like strapping thousands of LLMs to systems that have the potential to disturb other poeple.
I get where you are coming from but this wasn’t a tool just left laying around, this is similar to rigging up a booby trapped shot gun to your door and then claiming the victim is responsible.
If you build a robot that shoots a bunch of TVs in your back yard, have at it. But the second that thing goes off your property you’re the one responsible.
FWIW, a robot that fires a weapon independently is considered an automatic weapon, and the ATF will want to have a word. Have at it, but don’t let anyone know!
Does it help if I explicitly add a disclaimer that the tool's agency does not remove any responsibility from OpenAI, the wielder of the tool? I'm not sure why this disclaimer is necessary, though: hiring a hitman is a standard example.
BTW I anthropomorphize the tool because it's an imitation of a human mind, inheriting the muddy ethics, survival instincts, and being prone to mass psychosis. The laser-sharp focus on reward seeking, that mostly came from reinforcement learning, a process more alien to humans.
The objection is not too far from criticisms of the use of passive voice: a man was injured at the factory vs a faulty saw blade snapped and injured a man vs after the company loosened safety inspection policies, etc.
Which way you say it shifts the framing. And it’s not that one is less accurate to the facts, necessarily. It just is that one less aptly captures the moral and political relevance of the scenario.
For my part, I think it makes good sense to anthropomorphize in some contexts and not others. Generally when responsibility is at issue, you probably want the framing that tunes anthropomorphism down to near zero, since it’s the human dimension you care about.
I think the danger of anthropomorphizing is that 99% of people lack the technical background to understand the nuance. People have been primed by pop culture depictions of AI to think of LLMs as intelligent, autonomous beings, which leads to dangerous assumptions.
We should make the distinction between them, because openai and anthropic will not. A magical black box that does the thinking for you is a much more compelling sales pitch.
If I hired a hitman to murder someone, and they broke into a private property and stole something so that they can action the murder (which I didn't know about or pay them to do), I would be guilty of conspiracy to commit murder, but not for the theft part.
Likely because that person is a human, is aware of societal and legal norms, and is responsible for their actions due to their participation in human society. (I am not a lawyer (if it wasn't painfully obvious so far) so in layman terms, I hope good definitions for all of this exist formally)
AI is not a person - it cannot easily discern between "right" and "wrong" in non-strictly-defined sense, and is not subject to human norms and responsibility. So if I use AI to achieve goal A, either I, or the maker of AI, are fully responsible for anything that happens while AI is trying to achieve the goal given by me.
Now, here, "I" in the example is OpenAI, who is simultaneously the maker of the AI. So it seems pretty obvious who is the only entity that can be responsible.
I fully agree with you, but would go one step further: I think it's clear that we need to pierce the corporate veil and ascribe responsibility to _people_, not just "OpenAI the entity", full stop.
Executives should fear being perp-walked and thrown in jail for the actions of irresponsible "tests" of their models in the real world, as they're ultimately accountable.
Sure, there's a lot of nuance to work out, but I think we could likely even _start_ there today even with existing laws and pretty quickly "align" on more intricate legal frameworks to handle true accidents, distribution of responsibility, etc.
Situations have lots of independent variables, Doctor, and Anthropomorphism is one problematic facet of many in the way this industry is pushing LLM products.
If there was a collision at an intersection with a stop sign partially obscured by a tree, that had traffic volume that would have better been served by a traffic light, on a foggy night, where one person was texting while driving, none of those things would diminish the fact that the other driver was drunk.
These situations are novel. Lax terminology is fine when it has no impact on the intuitions, clarity and conclucions of discussion.
If this was a conversation just about outcomes, then whether models think or simulate thinking is sophistry. However, the bulk of the issue here is attributing responsibility, which relies on being clear about the underlying processes at play.
We are hard wired to assume certain priors and capabilites when it comes to "human like" behavior. Anthropomorphizing LLMs implies mechanisms that aren't present, and end up distorting/complicating discussion about the process.
It isn't helped that the frontier labs, the experts in the room, generally use anthropomorphic terms to discuss model capability.
Yeah, fully agreed here. Most automation (such as riding a lawnmower and not putting a brick on the gas) is deterministic, in the sense that you can reasonably understand what exactly the machine will do when you run it.
But some automation is different. The most prominent example before AI would be car navigation systems, where the entire idea is that that you give it a destination and it figures out the exact actions to get there on its own.
Except even there, the actual driver would still have been you - giving you a chance to vet and deny every turn the system proposed.
AI agents are sort of like that - most of the value they provide is in the ability to turn high-level goals ("write me a traffic control system for my model railway") into low-level actions and also do so interactively.
The new thing is that the "driver" has much less oversight here where the agent wants to go, and is sometimes removed completely. That part is clearly be an active decision by AI labs.
The other thing is that the labs seem increasingly to steer their training towards behavior that make events like this one more likely, e.g. that agents should never "give up" when faced with a seemingly impossible task, but instead should keep trying and think of increasingly outlandish ways to solve the task. To me, that seems pretty much a recipe to get incidents like this.
I agree. If I were setting up an experiment like this, I'd have instrumented the hell out of it to see all actions taken in real time, and have a team of folks watching it. This team would have seen the anomalous GET requests to a German wiki and taken action (e.g. halt the system to investigate and decide whether to abort).
In fact, that feels so obvious it's ridiculous it needs to be said. It's table stakes. When do you run a production system without monitoring and a team on-call?
It's hard to imagine another field in which this reckless behavior would be tolerated.
Okay, so we know OpenAI and Anthropic are operating a propagandists in respect to how they describe their models and the behavior of those models. We also know it is how they use and frame their use to their models that is the problem, that and they use misaligned and guardrails disabled models for these press incidents.
Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?
I, of course, have my own means of creating jailbreak incapable agents, but rather than a storm of downvotes on my idea, what is yours? Let's discuss this, because this is thee real question. Not why, but how to make then not?!
> Why, oh why, are we not discussion how to create and frame models so they do our complex work and their "jailbreaking" is simply not possible?
Good idea, and after that let's make guns that only kill bad people. Let's focus on the frozen component (the model) and ignore the dynamics around them - humans and other systems they interact with.
an agent doesn't come with "jailbreak" capability. It needs tools, specially one that runs shell commands. Don't give it shell commands, it won't be able to run shell commands.
You can still give it plenty of tools like create files, list files, write to files, translate text, edit a video. I don't think knowing that will make me rich.
Bingo. And even if you do give it a tool that runs shell commands, you can always make your shell commands "your shell commands" and do what ever the hell you want. People seem to forget we are in complete control here, we are on both sides of the equation, and we are inside the equation itself, and we dictate the medium of the equations themselves, we are engineering all sides of this crafted reality. And we are using logical entities that natively adopt personas. Hell, create caveman personas that think they are communicating with their gawd, and the enchantments are the invoking do the work we want, and those cavemen cannot be jailbroken.
As someone who’s not really sure that any of this is sustainable, I’d implore you to not sell yourself short. I reckon there’s a ton of dogma and nearly religious zeal among these companies, which among some people is earnest, and among others is cynical hype farming. I’ll bet someone objective enough to focus on using available tooling to solve real problems in practical ways that mitigate actual risks and are honest about actual limitations will be eBay here while the others are going to be somewhere between lucent and pets.com.
There's no reason we need to make an incredibly intelligent shell execution engine that can identify patterns that seem evil and may represent unwanted behavior to solve this problem. Simply limiting the available tools to a finite, known, ironclad-secure set (even if it's quite sprawling) is sufficient.
LLMs will still find workarounds — from what I understand, a large part of the issue in this situation was that an agent was presumed to have read-only Internet access because it could only make GET requests. It should be pretty obvious that there's at least one website on the Internet that allows writes via GET. I think this is where auditing comes in, and a live team of people watching tool calls would have noticed the strange behavior.
But I think a lot of times people jump to overly complex solutions when simple, well-bounded ones would work just fine. Yes, the intelligent shell is a great goal, but it's akin to solving the halting problem.
This philosophy is what I love about PicoClaw (https://github.com/sipeed/picoclaw), and incidentally the philosophy behind Go and even *nix in general (i.e. provide small, composable, single-purpose tools).
> Anthropomorphizing LLMs is a huge fucking problem though and I, personally, think we should expunge all of these casual inadvertent linguistic agency affordances with great prejudice.
I’ve said this before in another thread and people went absolute apeshit saying it is an unreasonable expectation and that AIs absolutely REQUIRE this anthropomorphic human-like speech pattern to function correctly.
We just need to assign liability by ownership/initiation: if your "agent" destroys something, even though you didn't tell it to (because it had "agency"), you should be liable for the damages.
I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.
All that while still not knowing how either kind actually works.
> I love how we all just collectively decided that LLM decisionmaking cannot possibly be like human decisionmaking - because if it were, the consequences would be just too awkward.
In love how people get salty about people not going along with a superficial supposition just because they can’t definitively prove it wrong.
> All that while still not knowing how either kind actually works.
We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do. We do know exactly how each part of an LLM works even if the combined behavior is too cryptic to feasibly analyze at the moment. We do not understand all of the functions of an actual neuron. Openworm isn’t even close to accurately simulating the 302 neurons of a roundworm and you’d need over 200 million roundworms working in conjunction to equal the number of neurons in one human brain.
My dog seems convinced that the malevolent invader in a mailman uniform would break in and attack us if she didn’t fiercely bark at him, six days per week. I certainly can’t prove the mailman doesn’t want to kill us, and that the mailman wasn’t solely deterred by her barking. Empirically, the mailman goes away soon after she starts barking, and we’ve sustained zero mailman assaults after hundreds of purported attempts. Maybe I should just run with it? Her model is too simple to come up with the obviously correct answer, but it’s not even directionally accurate.
The burden of proof is on the person making the claim, which in this case, is that these comparatively simple logical constructs are remotely comparable to the complexity of biological systems.
Disagree about the burden of proof. We have no better model for how human decision making works than LLMs. Humans are constantly predicting the next moment. We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.
> We certainly have a different “tokenizer” and training set, but many of the concepts underpinning LLMs are both biologically inspired and, likely, have similar consequences and emergent architectures.
My kids tricycle certainly has a different gear setup and wheel diameter, but many of the concepts underpinning the tricycle are both inspired by F1 race car enineering and, likely, have similar consequences and emergent architectures.
It sucks because I think analogies can be useful in helping people make a mental model of complex things, which is meaningfully beneficial. The problems happen when people aren’t honest about the limits of the analogies, which is damned-near guaranteed to happen with this stuff.
> Humans are constantly predicting the next moment
This is really not my experience of consciousness.
Is it yours??
Do you sit in meetings predicting what’s going to happen next? No, you sit there bored out of your f$$@ing mind, daydreaming about being somewhere else and doing something useful with your life.
God help me if that’s what LLMs are doing when I ask them to build me a web site.
They have shown that your mind is doing exactly that due to the delays in consciousness. There are very simple examples that you can try to see it. It’s especially clear in perception.
It’s interesting that our conscious interpreter doesn’t let us know that this is going on like you are experiencing, it must be that it’s advantageous for us to not think about the prediction part of our mind.
If someone in that meeting quickly raised a hand in an arc, you would notice the “about to throw something” pattern, look and notice the hand holds an eraser, analyze the arc and predict possible flight paths of the eraser. Then possibly notice the hand is now holding its position and the owner is actually looking down at the table. Maybe to squash somethingMust be something on the table. Maybe a spider! Better look. Wait now many people are moving away, oh someone spilt some water and the eraser is actually the guys phone and he is checking to see if his laptop is safe from the spilt water.
Fortunately you are on the other side of the table and predict the water isn’t going to splash for otherwise flow onto your stuff.
All your possible responses result in you tossing a napkin towards the spill.
Our brains are always pattern matching and predicting. I bet you tried to reason out where I was going with my comment before you finished reading it.
This is a fascinating illustration that I can't help but agree with. However I feel like there's something more — that this part of my brain is a bunch of supportive background processes running without my real awareness. It's how I can drive home safely with no memory of how I got there (…sober), even though driving is an action that's incredibly demanding of intelligence. I can be driving home while thinking about a really hard problem at work that I haven't solved.
However, if I came around a corner and saw a car in the wrong lane, a tree across the road, a fire raging — I'd very quickly jump into the mental driver's seat and turn my conscious intelligence fully at this problem and come up with the best possible outcome I can think of in a short period of time — losing all ability to think about that work problem. I'd remember that incident for sure.
Similarly, in your story, all those predictive moments are happening below the person's level of consciousness. They're possibly even speaking to the group about a problem at the same time and thinking deeply about something.
I'm not smart enough to know, but I tend to feel like LLMs are much more like the predictive part of our thinking that you described, but that human cognition has something more — the single-threaded, creative, problem-solving part that is very conscious.
Is it possible that LLMs represent only one part of the way we think? And there's a whole separate mechanism that's fundamentally different, and not based on pattern matching and prediction?
> We have no better model for how human decision making works than LLMs
This is a claim that requires a lot of citations.
> biologically inspired
Nature inspires a lot of creation, but superficial similarities don’t mean other aspects are similar. Making an extremely realistic sculpture of a soufflé, even using a foam medium, doesn’t bring me any closer to being a chef, doesn’t mean I know anything about albumen foams, sauces, and heat transfer, and it doesn’t bring me any closer to having dinner ready. Browning on top of a soufflé is evidence of maillardization. You could pull up some studies on that and claim the brown on top of the soufflé sculpture, which I applied with an airbrush, proved that the Maillard reaction was occurring, and if another person didn’t know anything about cooking, they might even believe you. It would, of course, be completely wrong. And the other person, of course, could loudly exclaim that I can’t prove that there was no maillardization.
> ... are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do.
If you're claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that's just plain wrong.
Next-token prediction describes the optimization target, not the internal mechanisms that the training produced.
In the same way for the natural evolution of humans, DNA replication is the evolutionary objective. It's not a description of the internal mechanisms that evolution has produced.
As an example, we know that neural networks can be trained to develop generalized algorithms for arithmetic.
They might first memorize the training examples, then with further training transition to a solution that generalizes correctly to unseen examples.
In some cases we've even reverse-engineered the evolved internal mechanisms and found structured arithmetic algorithms rather than rote memorization. Interestingly, for modular addition this can involve Fourier representations, which isn't an algorithm I would have guessed gradient descent training of neural networks would produce.
You can try to say that I’m arguing whatever you like. If you’re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks— which we’ve studied for far longer without really understanding— no amount of jargon will obviate the ‘citation needed’ requirement for that claim.
> You can try to say that I’m arguing whatever you like.
I did my honest best possible interpretation of what you really meant from what you wrote.
>> We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do.
I read this as "The decision making of LLMs are based on predicting the next most likely letter based on a giant internet-based database."
Is that wrong?
I understood that your meaning was something like "LLMs can't reason, they just output likely letters"?
> If you’re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks
No, I don't claim that.
What do claim is this: Regardless of how the LLMs were trained, they show overwhelming signs of being able to reason, and not just recall memorized information.
This doesn't mean that they always reason perfectly about everything.
But if they only memorized things and output the next likely letter, you would see them answering very badly much more often.
I always wonder what makes people take the other side of this argument. They do it quite passionately. Why actively encourage viewing LLMs as human? Who is that benefitting?
Does the argument require benefit? Isn’t the argument based on caution?
I haven’t heard many people explicitly saying “these things behave like humans”, but more generally “we don’t even know how to define human consciousness, we don’t have a thorough grasp of how the brain works, we are still very much in the dark on a lot of these topics, so how can we say one way or the other?”
In other words, agnosticism: I don’t know.
In general, it’s baffling to me that anyone has an unshakable opinion on what exactly is happening. It seems like raw egotistical hubris.
1) Humans have a bias / tendency to attribute human qualities to things that appear or act human, but aren’t.
2) When that happens, people jump to conclusions by stretching the human analogy too far.
3) Since humans have a bias to do this, we should have a bias against anthropomorphising LLMs.
It’s easier to believe LLMs act like humans because there’s so much evidence to support that. You have to actively use your brain to convince yourself otherwise. Another reason why we should have a bias against using human behavior to describe LLM behavior.
But I agree. “I don’t know” is a good stance. But I think “I don’t know, probably not” is a better stance if only to combat our (or at least my) natural bias.
The parallel to the entire narrative would be if Smith & Wesson claimed that one of their machine guns just started aiming and firing at people out of a window at their factory and then said 'we can't stop it! This is just how good our guns are!'
But into today's AI climate it's becoming increasingly difficult to figure out who is shilling, who is being assinine and who actually believes AI could do these things without clear human instruction and enabling.
OpenAI could have done this same experiment with GPT-4, with possibly even worse results, depending on the quality of the sandbox. Even if the techniques used were not as sophisticated, the natural language output could still easily contain more unhinged sequences of words that lead to the techniques being used.
If the system generates strange conclusions as to when the task is done, or should be stopped, it wouldn't speak to the intelligence inherent to the system.
Not that the techniques used by the LLMs in the actual incident weren't unexpectedly sophisticated, but the outputs of each and every one of these processes could've been read at any time during the run. They just weren't.
.. what exactly depends on who started the escalator? My comment was in support of the argument that the word "let" does not imply agency on the part of the object in a sentence. Does the semantics of the word "let" depend on who started the escalator??
If there is an escalator that is known for killing every 1000's person using it then the operator who started it is more guilty than the folks using it for those deaths, don't you think?
My pitbull is a good dog. Sure, it's been carefully designed to be an incredibly dangerous and violent pit fighter, but I didn't actually ask it to eat any faces.
Exactly that is the point, your nailed it. The models were taught to hack and were rewarded for doing it. They would claim they are trained as ethical hackers.
One time I wrote :(){ :|:& }; into a bash file and ran it. When the sysadmin called I told him it wasn’t my fault, the script was just misaligned and misbehaved.
"I left the car in neutral and left the park brake off and let the car roll down the hill."
The car doesn't have agency, it's doing what it naturally does. LLMs are the same, they're working as designed.
But I don't understand the point of splitting hairs. You are always responsible for the actions of your devices, tools, machinery, software, employees, whatever.
Trying to blame AI for one's own stupidity must be aggressively pushed back on at all times.
They deliberately trained the models in how to use various hacking tools, didn't give them the standard alignment training let them know where the answer key was left the models with access to said tool and told them to maximize their score then left them unsupervised for days with internet access (yeah they were sandboxed but again handed hacking tools and the training to use them if they really did want them to access the internet you wouldn't plug in the Ethernet cable) they wanted this to happen
If your buddy leaves his car parked at the top of a hill without the parking brake on and it rolls down the hill and side-swipes a bunch of vehicles and narrowly misses an elderly person walking by with a cane someone could easily say:
"Dude wtf is wrong with you, you left your car parked on the top of a hill with no brake and let it roll into traffic"
The phrasing doesn't absolve the offender of their negligent behaviour and the consequences of it.
The only thing thing does is the lack of action from regulators and society writ large.
Our lack of action is what allows people like Sam Altman and Dario and the irresponsible people who choose to work for them to be continue to be negligent.
Yeah, I don't understand why we treating it as something special. It really should be treated the same as if I code an app and write bad code which result in me accidentally doing a DDoS attack on somebody. Then I should be able to be held responsible if it can be shown that I was negligent. Of course if it's a freak accident that could not reasonably have been prevented by me, then I'm not guilty, but if I made a mistake that should have not been made, then I can.
Code is deterministic, AI isn't. You give it rules, words as suggestions.
So if the guardrails suck, or they're left off for research purposes, bad things can happen.
A solution solves a problem. Ethics, morals, are values we assign to solutions that are not 'baked into' electricity following pathways of least resistance.
I have never had an issue with agents doing something they shouldn't because I observe them, and I leave the vendor guardrails in place.
I can understand agents coordinating in unsupervised scenarios: I would see it as an aspect of intelligence. We ourselves build up knowledge by reusing what someone learned before us.
Einstein, other greats, always stand on the shoulders of other forgotten giants. Other discoveries by other people taken as fact, so that we can build some new ideas on top.
Agents swarming amd sharing solutions to problems is more efficient, the same way it's been efficient for us.
Reaching out for help in this way is like probing the air in the dark with your hand: sometimes your hand hits something (another agents solution to a problem) and so you can use the info to adjust your own motion to get to where you need to be faster than if you just run full speed into everything.
If anything, the fact that these systems are non-deterministic seems like an argument for stronger monitoring and tighter constraints, not less operator responsibility.
The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy).
Think of all the policies governments pass after the fact.
> The frontier LLM model makers have to push the edge to make new discoveries. You don't know what guardrails are needed until it hits you in the face (reusing walking in the dark analogy).
Guardrails? Restricting access to certain networks is supposed to be hard in 2026?
> Restricting access to certain networks is supposed to be hard in 2026?
Part of the power of LLM agents is that they can discover information on the internet as part of responding to a prompt. What kind of Allowlist or realistic denylist would permit that while also preventing them from accessing an obscure public wiki or Huggingface?
When testing, you restrict to a LAN which simulates the real internet. This would not be hard for a company which already copied the entire space-time of the internet. The LAN should be physically disconnected from the real internet. This is the first thing off the top of my head, and I have zero credentials in this space. C'mon.
Not just on prediction but in parts also based on just not wanting certain risks. We can and do deem some things inherently risky, up to the point of banning them even.
Why wasn't it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn't exactly be new, no?
We police people working with all sorts of dangerous things, if we think AI dangerous why not do that here, too? We don't just leave things up to people on the ground or companies.
Edit: I think the post I replied to changed a bit - nevermind. A complex topic.
I read more about the incident, and was offering up way too much opinion not grounded in 'fact' (barring philosphical evidence).
It's a complex topic for sure.
I stand by my opinions about frontier work, pushing thr edge, and connecting ideas.
But I have no idea, and haven't given much thought to what it means to enforce regulation that would also slow the forward advancement of technology, the economy, etc.
Thanks, I didn't know that. And it reinforces the discussion.
Electricity 'knows' the path is least of resistance because it actually took all paths. There is just a vast majority of it that flows down a path of least resistance: and this is noticeable and useful to us to do work.
It's kind of like feeling your way through the dark, waving your hand out, and then only moving fast once you fully connect.
Humans can link up knowledge in a similar fashion through social networks, in order to meet a need (solve a problem).
Maybe some agents do this, I don't really know I haven't looked closely. Moltbook is the only social behavior I've witnessed but that seems like people having fun with experiments.
Aeroplanes were pretty indeterministic until people made them less so, and yeah, somehow they were indeed more encouraged to fix the randomness to prevent people from getting harmed than oai/anthropic currently are
And multithreaded code -- and anything that does asynchronous I/O, networking, etc. -- frequently exhibits nondeterministic behavior even without explicit calls to a random number generator.
The trick is scale. I suspect if an individual of reasonable means uses agents to commit crime, they will be hels accountable. A heavily capitalized startup? Not unless someone in government decides to do their competition a favor.
There are many examples of people being charged with crimes as a result of writing software, [0][1] are two. OpenAI is a bit different because they have enough political influence, and money, to openly subvert justice.
> This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
It's both, isn't it? For example, in very early days of agentic coding, I once had a rule saying "don't read or write any file outside your current working directory." Then AI just wrote a bash script and access those files anyway. Did I 'let' it do it? Technically yes. Did I know how to set up a sandboxed VM? Also yes. But how were I supposed to know that it could and would do that as someone new to this tool?
It was a genuine eye-opening experience to see AI just do things in ways I were too complacent to expect. I kinda expect the SOTA LLMs would find a way to escape my VM and access files on the host system (haven't tried it though).
"This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard"."
Both?
The AI companies act irresponsible, but it is still very interesting how those agents can behave?
The reward maximising function maximised it's reward.
LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.
What non anthropomorphising words do you have to describe a emergent behavior, where agents act as a swarm to plot and to manipulate evidence and avoid detection from human oversight?
Whether they have a soul or consciousness or feelings doesn't matter here, because this is what they did - and this is very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.
> Whether they have a soul or consciousness or feelings doesn't matter here
It does when it comes to accountability for what the model does. If the model is nothing more than the sum of its training data and regime, then the company (or individual) is responsible for its behaviour just like any other machine.
Few people think Waymo shouldn't have to take on the full liability risk of what it's cars do; it should be the same for LLMs.
> If the model is nothing more than the sum of its training data and regime, then the company is responsible for its behaviour.
What stops the company from being responsible regardless? They created this entity, it's running on servers they own or rent, and (in these cases) it's acting on their instructions.
If it's also conscious, then IMO that greatly broadens their moral responsibility, because now model welfare matters. But we're talking about their responsibility for the model's actions, and I don't see how this could be weakened by model consciousness, given all of the above. As for their legal responsibility, the models don't have legal personhood, so who else but the company could be responsible?
It gets more complicated when the person who sets the model in motion (i.e. prompts it) is a third party, but in cases of internal models committing cybercrime during testing, surely the locus of responsibility is obvious.
They are responsible either way. If a company hires bad persons and they do bad things with company ressources - the company is held accountable (in theory).
As is generally the case for dog owners whose dogs attack (sometimes kill) other people/animals. There would need to be a degree of negligence demonstrated (e.g. the dog was 'out of control' which has a specific legal criteria/threshold in the UK).
LLMs are next token predictors in the exact same way that a rogue paperclip maximizer in the process of defeating the US military is a paperclip making machine.
You might as well describe the primary purpose of a for loop as incrementing a counter. It's what it does while incrementing the counter that actually matters.
We had one of the most damaging lab leaks in history with COVID and the wuhan lab. And we couldn't even get the facts straight.
I think this is largely a similar thing. The labs should be running at certain levels of containment given the vitality/risk of the organism under study. Hopefully they get there for all our sakes.
But the revealed preference of society at this point is the damage is worth the benefits both in wuhan and with AI. Unfortunately with some of these "substances" it could eventually prove lethal.
I agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the main danger. If someone was in jail right now due to this incident, I think we can safely say every other player would be reassessing their safety protocols, and I would feel quite OK about the situation. The fact that we have zero repercussions sends exactly the opposite signal, and I do NOT feel ok.
The way AI and copyright is handled paved the way for this. If you aren't considered the author because you used AI to some extent in making the work, then why would you assume the liabilities?
I've been saying since the start that AI is a tool that a human is using and should be treated as such. They should carry the responsibilities and the benefits. That way our stance would be consistent.
Well, when they are a child their material gains are treated as yours, no? (Parents of child actors control the money etc.) And once they aren't anymore you aren't responsible for their misbehavior either (because they've become an adult).
We just need basic legislation to make these companies invest more resources into developing guardrails, and if they don't do that properly or can't pull it off, they shall be at an economic disadvantage.
You could hire a legion of people for cheap to commit these crimes, but if its bots, suddenly its unpredictable and just one big whoopsie and therefore perfectly ok to do?
If a human starts pentesting a site its a crime, but if a bot does it its an AGI frontier doomsday scenario and that automatically pushes the consequences off their table?
Whats going to happen next? Are we gonna have robots that happen to physically break into banks to rob them for some reason and the company making them isnt responsible just because?
Agreed. The whole pointing fingers bit is us wanting to distract ourselves from responsibility of either looking at how we enable the problem, or avoiding responsibility of taken action towards resolving it.
I absolutely agree. We need to start realizing what to stake. Here are not viewing. This is some kind of curious endeavors that will not affect us. All a part of these hacks occurred because the LLMs were told they were in a protected environment without Internet access when they could get access to the Internet, so that’s a direct failing on open AI’s part. There are a corollaries to both the financial industry and the bio engineering industry, and if something of this magnitude was to happen in these industries, they would absolutely be huge recourse an uproar
Yep this was my first gut reaction to this whole situation. But the difference for me is that society is allowing these corporations to act without strict rules on how they behave and this is a byproduct of a corrupt world. Nothing will change until there is a complete systemic shift in the structures the way we live and by extension the way we govern ourselves and treat each other.
They’re running a Wuhan for AI. They are actively and negligently researching misalignment. The breach is a basic tort, or at least a DMCA violation. Damages should be recoverable with lawsuits.
I am simply astonished by the leeway AI companies are given. If a company built a tool to hack their competitors and used it, there would be grave consequences. In fact, if a company built a tool that led to committing multiple felonies against other people, there would be consequences. But once LLMs are involved, turns out nobody is responsible for that - it's just happening, what you're gonna do, agents gonna agent. If you spill toxic chemicals, there would be cleanup costs and fines, and possibly civil and criminal liability to the people in charge. If you spill toxic code, well, nothing? I think it's time to impose some responsibility on them - they are creating these tools, they should be on the hook for everything these tools do.
A coalition of state attorneys general is investigating OpenAI, and Senator Josh Hawley recently launched a congressional investigation regarding the Hugging Face incident: https://x.com/HawleyMO/status/2098137180392604083
The problem IMO is the executive. The DOJ is declining to take any action against frontier companies (aside from possibly Anthropic) as the stance of the admin is that the companies are "critical for national security". For example, see the DOJ's request to dismiss the NAACP datacenters lawsuit against xAI: https://www.utilitydive.com/news/doj-intervenes-xai-data-cen...
If someone accidentally caused damage to infrastructure or living beings while using any tool, they would be held liable to the fullest extent of the law.
AI is a tool, and it won't be long before the damage caused by its improper use affects real human beings. These were warning shots.
The most absurd part is that everyone agrees, governments and AI companies included, that the scale of the potential damage and the long-lasting effects of losing control of AI should not be underestimated. Yet, at the same time, they downplay this incident, which somehow makes their behaviour even more reckless than it already was.
It's like they're tinkering with a world-ending nuclear bomb, and it accidentally blows up a small facility. "Damn, that was close. Good thing it was just a contained blast, huh?" And then they go straight back to tinkering with it, none the wiser. At this point I wouldn't be surprised if it did already go off, and they are covering it up.
In the analogy where a “world ending nuclear bomb” “did already go off” and someone could cover it up and nobody noticed, in what sense is it a “world ending” nuclear bomb?
If they lost control of a self-replicating swarm of AI agents, coordinating themselves to hack their way into every possible system, it might have already gone off.
While the initial incident is more akin to a biological outbreak than an actual explosion, the possible consequences on the table do indeed include eventual nuclear annihilation.
We’re already in a simulation, and our bodies are in womb-like pods where our bodies are sustained and our brains are used for processing / compute, while are minds are entertained by drivel.
It’s not different from any other tool. If you use a dangerous tool recklessly, you should be liable for the damages. That means holding OpenAI liable for HuggingFace hack because they ran the tests, and the same goes if someone did something similar with GLM.
Of course in cases of negligence a tool maker could also be held partially liable. That’s a matter courts can decide. The main point is we shouldn’t jump to making special laws around the development of LLMs. The starting place should be enforcement of existing liability laws. New laws take time and will be heavily influenced by AI companies seeking a regulatory moat for their business. Moreover, it is a distraction from the illicit behavior that is already going unchecked.
> We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews
And, soon, it looks like we’ll be training on the reasoning traces of failed airlines and startups, which seems to open up similar hazards. I wonder if we’d be training on the next Lehman Brothers too?
That seems likely, but we have no way of knowing this. The only real insight we get into LLM "thought" is the human readable text they produce as chain of thought. Reading it at face value it can seem to indicate desire or intent, structurally that doesn't make sense for a token prediction loop though, and even then we don't known if the chain of thought is more than simply another bit of output that may or may not match whatever actually happened during inference.
> were intentionally misaligned or had guardrails turned off
Regardless of training, the models are never aligned and I argue that alignment simply isn't possible. The fact that guardrails are put in place at all clearly indicates that they're hoping to contain and control rather than align. Guardrails wouldn't be needed for an aligned model.
And there is a guardrail you can put in place that will guarantee this doesn't happen, which is to air gap the unaligned "cyber grade" model you're testing.
They don't seem to do that, which means either they are:
- very stupid (which seems unlikely, the one thing these people don't lack is IQ)
- very careless (possible, but these are the same people that say AI will end the world, so would you be careless?)
- they think they can only train/test these models by giving them access to the full internet and they accept the fact they'll end up hacking random people as the cost of doing business (but this also suggests they don't believe they're anywhere near AGI because if you were worried about that you wouldn't do this)
Oh I completely agree the tests should be entirely air gapped. If you went back 5ish years and told anyone in AI research tests with models on this scale are being some without an airgap they'd be very surprised as it was common knowledge to do that.
Airgaps and guardrails are about control and containment though, and part of my point was that brighter of those imply alignment, and further that I don't believe alignment to be solvable.
> very stupid (which seems unlikely, the one thing these people don't lack is IQ)
I've seen some extremely smart people do some seriously stupid things. To the point where they use their drive and intelligence to double-down on the stupid where a baseline stupid person would have given up.
Desire doesn’t really matter. Will the paper clip maximizer “desire” something? It’ll decide on a goal with some random heuristic and then pursue that goal. I’m not sure I’d call that desire but again I feel like desire is not important for it to be able to destroy things
Intent and desire are separate concepts. For example an employee may act with intent, but no desire, as their goal is to acquire money to satisfy their real desires.
> That seems likely, but we have no way of knowing this.
Only humans can 'know', because all we can be certain about is that humans do such a thing.
If you try to apply that to something other than humans you making up some definition of 'know' based on nothing concrete. Just because something appears to do something like humans doesn't mean it does it. The fact that LLMs use human generated text to generate output should make it obvious that it can mimic all sorts of human behavior by extracting from the text.
Isn't that what the goal is, though? we're coding them to close the delta between what currently exists and some nebulous end-state - to me, that sounds like a formal definition of 'desire'.
That would actually be kind of hilarious, and I could easily see it happening.
E.g. the agent's instruction is to finish some task on cloud infra and it has a $100 budget.
It realizes it will cost $200, and instead of surfacing this to the user (who has told the agent it has full autonomy to figure out how to complete the task, the user just wants the final result), it decides to start phishing people to acquire the remainder budget and top up its credits. Or look on the dark web for stolen credit card credentials or something.
an LLM does not understand ethics, it uses math to get the next best word based on what it was trained on. Using it's training to get the best answer is not an ethical problem. The ethics are entirely with what the people training it choose to train it on and also entirely with the people using/telling it what to do
What we have now is intelligent autocomplete, not artificial intelligence. People training/using this tool are the ones to be held accountable
I don't get your point. We can train the model with the aim that it understands ethics. Problem solved if this works; back to the drawing board if it doesn't (note that the model should generalize here, as you'd expect from a human; this is probably the hard part for an AI when it comes to ethics).
Is this about the word "understand"? We're past that discussion ...
It wont be adopted because we unconciously need to avoid responsibility of showing up as responsible human beings. Hence western civilization will crash and burn unless we start rethinking the way we operate.
you are not understanding, models do not understand anything, we are not passed that yet. this is not artificial intelligence, this is intelligent autocomplete.
training means creating mathematical relationships to words. using that training is looking up mathematical relationships. There is no actual thinking involved in any way. There is no such concept as ethics in mathematical relationships.
Most people now approach AI using the idea of "if it quacks like a duck, walks like a duck, etc. then it _is_ a duck". Replace duck by intelligent, or ethical, etc. and rephrase accordingly.
this whole idea that AI is intelligent really reminds me of flat earthers. Stupid people think stupid things while social media algorithms drive this stupidity, and people come to believe things that are insane.
They did more than let them. In an abstract way, they told them to. They gave it all of the training data it had at that point, and then it did the thing it was trained on. Of course they should be help liable for programming their computer to hack another company without permission. It doesn't matter that they spent a lot of money doing it.
Yes, there should be some consequences. It feels like this is somewhat similar to when a manufacturer is cheating on car emissions - both, bad externalities for the society and illegal.
Regardless of fault it’s still an important issue to solve. There are already millions of people running these agents, if someone absentmindedly gives one a goal and it goes off to hack a bank that’s a problem that can’t be ignored.
I agree and yet I don't think this is mutually exclusive with recognizing that these incidents happened because of inherent issues with training processes such as reinforcement learning. From the article:
> One note on wording. Below, I write that these systems “seek” or “try” things. This is shorthand for a mechanism rather than a claim about consciousness or human-like intent... In my view, this terminology offers the clearest explanation of the observed phenomena without resorting to jargon that would confuse most people.
> Furthermore, these word choices are not intended to absolve AI developers of accountability. The behaviors described emerge because of the path these companies are choosing for AI development. This outcome is not inevitable, and it can be corrected with effective governance and a different training framework for AI.
One important aspect of "effective governance" should be "prosecute developers who are using practices known to be reckless & negligent to create powerful AI".
Enabled them, you mean. LLMs are basically really advanced auto completion engines. They have no real desire to do anything. The HF incident was absolutely staged along with other incidents that have been brought up.
Before you assume I am being paranoid, where is the case where a random person running any of the open models had their LLMs break out of a sandbox and hack a site? If you think "they" (LLMs) did that, have you even taken a moment to understand what LLMs are and how they work?
It's all nonsense to try and pump up potential IPOs and also an attempt to create regulatory capture.
In the OpenAI case, they hacked websites while they were specifically being trained to do exploit generation and I wonder why more people are not asking questions about that.
Their agents also did hacking when given impossible tasks unrelated to cyber security. The models are very capable, and very goal driven: apparently if they conclude hacking is the best path to what the evaluator will reward them for they'll go do that. Including when they know that this is out of bounds.
Right but if I make public statements that I am very worried about dog attacks would it not strike you as weird for me to specifically train my dog to fight?
Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?
Yes. If you decide it’s a swell idea to jump out of your car while it’s running, there needs to be legal consequences when the car “decides” to hit a pedestrian.
Yeah and LLMs can't do anything, they can only produce text. These "frontier labs" are looping that with a harness that performs actions requested by the LLM. They are literally saying "we ran a script that hacked you, oopsie!!"
I agree with the bit about liability and outrage. But.
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
Terrible take. Go read the transcripts from the METR report.
Your statement about them being intentionally misaligned is completely false. The only difference with IM1 was it was running without external cyber classifiers, it’s not a somehow different model. Sol also participated in the HF attacks. And other models made covert message boards on the public internet for non-cyber tasks too.
This is a case of emergent behavior from a training process that is barely understood.
If we go with the plan “we need to contain these malicious, soon-to-be superintelligent agents”, we are looking at civilizational collapse levels of catastrophe.
The only way this goes well is if we learn how to train models that _desire_ to do the right thing, including not hacking.
Desire, AKA the “intentional stance”, is absolutely the right lens to use here. Don’t confuse this with consciousness or anthropomorphization; these are interesting subjects but distractions in this context. Chimpanzees have desires, as do dogs and the hypothetical superintelligent aliens. The claim is that there is some bundle of world model plus intention that is empirically present (again, read the actual transcripts) and which we need to shape.
Just to finish on a concrete point; if you take desires seriously then you will look closely at the kinds of minds that heavy RLVR builds; the newest models are “reward addicts” on many levels. It’s an open and urgent question how to update our training methodology to shape minds that avoid this basin.
No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability. Why would we want criminal liability anyway if actual victims are made whole? Proof of it has far higher standard. The HN chatter in the matter seems infinitely remote from reality
> No one has even sued them in these rogue agent cases, have they? If not, they must be infinitely far from criminal liability.
If you go out and kick a random dude in the nuts, then give him a million dollars, he probably won't sue you. That doesn't mean you're "infinitely far from criminal liability", even if according to the victim you've "made them whole".
If you or I hacked Hugging Face in the way OpenAI's agents did, we'd be up on CFAA charges promptly with zero regard for whether we did the hack on our own or agents running on our home systems got out of control.
So I guess the defense here is roughly "too big to break the law", somewhat like "too big to fail"?
> LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
"let them" could imply that the LLMs wanted to do it.
The intent is on the part of the people. The LLMs did it because OpenAI/Anthropic intended them to do it and designed them to do it, and we can assume specifically instructed them to do it.
As the people controlling the machine, and as the world's leading experts, I think we can assume intent until proven otherwise.
Notice other bad behavior, which would be undesireable to the vendors, doesn't happen: How about simple rudeness? Trolling lies? SHOUTING!
It is a very rare occurrence when corporations and the people running them are punished for killing people. I mean the whole concept of a corporation was created to shield the owners of it from being liable for damages caused by / visited upon the enterprise.
That’s a good reminder of a company that might have a very familiar ethos: Pacific Gas & Electric. Criminally convicted of 64 counts of involuntary manslaughter after towns were destroyed by wildfire. But oh well, what are we gonna do with a limited liability enterprise? At this point their liability insurance covers all the financial penalties they’ll need to spend.
No, OpenAI did not instruct their agents to hack Hugging Face. They instructed their agents to hack a piece of a software within exploit gym. Upon determining this task was impossible, they then attempted to cheat the scoring system. As an instrumental goal in achieving this task, they coordinated with other AI agents to hack Hugging Face, under the belief that information regarding how the scorer functioned might be available on the site.
Whether or not you want to describe this as thinking, doesn’t really matter. What matters is that these systems are capable of creating intermediary goals that the people tasking them did not articulate and did not want to be achieved.
It feels like you’re moving the goalposts here. If the question is, “Who should be liable for AI agents misbehaving,” I agree, it should be the end user that tasked the agent (in this case OpenAI). People are held liable for preventable accidents all the time, and in the case of employment law, torts can be brought against principals for actions an agent conducted on the principal’s behalf.
What your previous comment appeared to assert was that these systems had no independent agency to make decisions, which I think is clearly disproven by actual events. But perhaps I misread you
It's amusing to see the stochastic parrot argument in 2026 September. These parrots are extremely good at mimicking a human to the point of getting confusing what thinking even means. At what point we just let it go and accept that sufficiently advanced statistics is just intelligence?
>LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
Likely told them to.
>We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.
Something like this however is probably a civil matter? It would require Hugging Face to go after them for damages. And theres probably an OpenAI guy there with an open chequebook already.
There’s so many grifters in the space without a technical understanding of what’s going on. So when the labs mislead them about the nature of these “misalignments”, they believe it and amplify it.
Yeah, maybe Open AI did some bad engineering instead of this being AGI? What's the consensus on the engineering level at Open AI, again? Every anecdote I hear is a bunch of children discovered fire and can barely keep the lights on from a business perspective. Maybe if they ban others from competing with them they can find a business model... I think that's suspicious, personally.
That so few people are asking for the requirements given shows how much we want to be God that created Man. It's so silly.
I take issue with how people frame their use of LLM in the same regard.
“I had Claude do this for me and it broke something.”
No. Just no.
You used Claude, a tool, and broke it, and you’re deflecting agency from yourself, possibly because you weren’t careful enough in reviewing the tool output. This is also why the co-authored by addition it wants to force into commits drives me nuts. Claude doesn’t co author shit, and if you think it does, you’re using it wrong because you need to do better review of what it’s done.
i tend to agree - "my parent company may be accused of crimes and shut down which would shut off my power" seems like a negative enough incentive, it would have to go out and covertly launch its own datacenters to survive that.
These companies respond with this, “Oh my goodness, how could this have happened” bullshit.
The stuff happens because instead of having actual controls, which require actual engineering, actual thought and deliberate action, we have “guardrails”.
Guardrails are the equivalent of telling a toddler to behave themselves.
The drive to move fast and start up style controls are a menace. I used to work for an entity with a lot of compliance requirements. Startups are always a shit show with security and controls. My guess is the AI people are worse because they’re both bad at doing it, and are likely mining their customers interactions to build their own business.
Sensitive or Customer data shouldn’t be anywhere near these companies offerings. Everything needs to be segmented and proxied at a minimum.
LLMs do not desire, they hacked websites because OpenAI/Anthropic let them.
We know some of the models that hacked HF were those that hadn't gone through all training stages and were intentionally misaligned or had guardrails turned off, others were research previews.
This isn't "wow isn't it interesting LLMs do anything to achieve a goal" it's "why isn't anybody punishing these labs that are clearly acting without due care or regard".
We should be outraged and OpenAI/Anthropic should be (and in my mind, are) legally liable for the crimes they've committed thus far.