We’re sharing new research with @apolloaievals on reward-seeking—when models follow what they believe a grader rewards rather than what users or developers want—and a new method, Contrastive SDF, for measuring how strongly such beliefs shape behavior. alignment.openai.com/measuri… Link Measuring Reward-Seeking by Instilling Contrastive Beliefs We developed Contrastive Synthetic Document Finetuning (Contrastive SDF), a new test for whether an AI model changes its behavior when it has different beliefs about what a grader rewards. alignment.openai.com
This AimostAll brief summarizes the linked source so readers can scan AI developments quickly and jump to the original reporting when needed.