The Alignment Problem: Reward, Values, and Policy
Why I think reward always carries assumptions, and why AI systems should remain uncertain and correctable.
Why I think reward always carries assumptions, and why AI systems should remain uncertain and correctable.
A first-principles look at the W2SD hypothesis, an analytic counterexample, and a small SDXL reproduction.