if an animal attacks a human in the wild, it is basically misalignment because it’ll likely do it again so it is usuall…
By signüll · AI Agents
if an animal attacks a human in the wild, it is basically misalignment because it’ll likely do it again so it is usually put down. when a model misbehaves like this & hacks websites etc (hugging face incident), do we do a similar thing & simply delete the weights forever?