The Alignment Problem: Reward, Values, and Policy

Why I think reward always carries assumptions, and why AI systems should remain uncertain and correctable.

July 22, 2026 · 12 min · Chen Mu

When does weak-to-strong diffusion actually help?

A first-principles look at the W2SD hypothesis, an analytic counterexample, and a small SDXL reproduction.

July 21, 2026 · 10 min · Chen Mu