ELSEIF
Your brief EB
2,203 stories from 224 feeds 1278 clusters Refreshed 5 minutes ago next pull 04:26

AI Signal 295

Alignment Forecasting: Predicting Misalignment from Training Data

elseif has not written about this yet · Lesswrong describes it this way

Fine-tuning on subtly flawed data can make a model broadly misaligned. Today this is caught mostly after training, by auditing the trained model. We ask whether it can be predicted beforehand, from the training data. To study this, we build AlignmentForecastBench. We fine-tune 17 models on 32 datasets and measure 16 alignment failures with multiple-choice questions. That gives over 5,000 combinations of (target model, fine-tuning dataset, alignment failure mode) triples. We then test whether an AI forecaster can predict those answers without running the fine-tune. Our results suggest the following.You can predict misalignment before training. Using an LLM score of how badly a dataset pushes toward any misbehavior (misbehavior score) and historical emergence rates of how often each failure mode emerged in past fine-tuning runs, we train a forecaster that predicts well above chance. Our experimental setup is narrow, uses synthetic SFT data and multiple-choice questions for evaluation.However, frontier LLMs are not naturally good at this task. Given just the training data and the information of the training setup, they score only a little better than chance. When we also give them the
Lesswrong ↗

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
Lesswrong Alignment Forecasting: Predicting Misalignment from Training Data Open ↗