r/MachineLearning 🤖 Ai 👁 0

Observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting? [D]

I've been running informal experiments on RLHF-aligned LLMs and consistently observing something I can't fully explain. Posting here to get feedback and find out if this is a known phenomenon or if my methodology is flaw

📄

This source provides headlines only. Use the button below to read the complete article on the original site.

📰 Read the original article on r/MachineLearning

Originally published by r/MachineLearning. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.