Blogs
If I have seen further, it is by standing on the shoulders of giants.
From SFT to RL: Reward and Policy Gradient
Where RLHF's reward signal comes from, and how policy gradient turns a sequence-level score into token-level updates.
From SFT to RL: The Two Degrees of Freedom
Why SFT is reference-distribution fitting, and how changing token weights and sampling turns it into RL.