Blogs

Blogs

If I have seen further, it is by standing on the shoulders of giants.

From SFT to RL: Reward and Policy Gradient

Where RLHF's reward signal comes from, and how policy gradient turns a sequence-level score into token-level updates.

18 min read

From SFT to RL: The Two Degrees of Freedom

Why SFT is reference-distribution fitting, and how changing token weights and sampling turns it into RL.

11 min read