Observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting? [D]
I've been running informal experiments on RLHF-aligned LLMs and consistently observing something I can't fully explain. Posting here to get feedback and find out if this is a known phenomenon or if my methodology is flaw
📄
This source provides headlines only. Use the button below to read the complete article on the original site.
📰 Read the original article on r/MachineLearning
Originally published by r/MachineLearning. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.