Take a model fine-tuned into misalignment and add a little of the word "safety" to every token it reads and writes. Does it behave any better?