We post-trained a model for offensive security instead of teaching it to refuse
Everyone's training AI to refuse. We trained a model to break in. Most other tools are just wrappers and the problem is that they still have all the guardrails on from the foundation model where they will inherit its re
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/netsec
Originally published by r/netsec. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.