r/MachineLearning πŸ€– Ai πŸ‘ 0

Context-Induced Activation Drift: Long benign context passively decouples RLHF alignment without adversarial prompts (Mechanistic Interpretability + Ablation) [D]

TL;DR: We observed that feeding a long, benign, thematically coherent context prefix ($L \in [100, 3000]$ tokens) into google/gemma-3-1b-it causes a massive passive shift in internal activations ($\Delta h_2 \approx 343

πŸ“„

This source provides headlines only. Use the button below to read the complete article on the original site.

πŸ“° Read the original article on r/MachineLearning

Originally published by r/MachineLearning. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.