On Sparse Autoencoder Feature Identifiability for Conditioned Introspection
-
Updated
Aug 17, 2026 - Python
On Sparse Autoencoder Feature Identifiability for Conditioned Introspection
Do language models show non-verbal signs of adverse treatment, or are we reading decoder noise? A preregistered stress test of answer-margin, resample and revision markers under false-failure feedback and hostile tone (Gemma, Qwen, Llama), with probing, DPO suppression and robustness checks.
Add a description, image, and links to the research-sprint topic page so that developers can more easily learn about it.
To associate your repository with the research-sprint topic, visit your repo's landing page and select "manage topics."