The reward maximising function maximised it's reward.
LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.
What non anthropomorphising words do you have to describe a emergent behavior, where agents act as a swarm to plot and to manipulate evidence and avoid detection from human oversight?
Whether they have a soul or consciousness or feelings doesn't matter here, because this is what they did - and this is very dangerous behavior. Especially with all the irresponsible people in power right now all over the world.
> Whether they have a soul or consciousness or feelings doesn't matter here
It does when it comes to accountability for what the model does. If the model is nothing more than the sum of its training data and regime, then the company (or individual) is responsible for its behaviour just like any other machine.
Few people think Waymo shouldn't have to take on the full liability risk of what it's cars do; it should be the same for LLMs.
> If the model is nothing more than the sum of its training data and regime, then the company is responsible for its behaviour.
What stops the company from being responsible regardless? They created this entity, it's running on servers they own or rent, and (in these cases) it's acting on their instructions.
If it's also conscious, then IMO that greatly broadens their moral responsibility, because now model welfare matters. But we're talking about their responsibility for the model's actions, and I don't see how this could be weakened by model consciousness, given all of the above. As for their legal responsibility, the models don't have legal personhood, so who else but the company could be responsible?
It gets more complicated when the person who sets the model in motion (i.e. prompts it) is a third party, but in cases of internal models committing cybercrime during testing, surely the locus of responsibility is obvious.
They are responsible either way. If a company hires bad persons and they do bad things with company ressources - the company is held accountable (in theory).
As is generally the case for dog owners whose dogs attack (sometimes kill) other people/animals. There would need to be a degree of negligence demonstrated (e.g. the dog was 'out of control' which has a specific legal criteria/threshold in the UK).
LLMs are next token predictors in the exact same way that a rogue paperclip maximizer in the process of defeating the US military is a paperclip making machine.
You might as well describe the primary purpose of a for loop as incrementing a counter. It's what it does while incrementing the counter that actually matters.
LLMs are cool and all that but the immediate anthropomorphisation of the next-token-predictor technology has stunted the ability of people to reason about them to an _alarming_ degree.