Tags


LLM

From SFT to RL: Reward and Policy Gradient

Where RLHF's reward signal comes from, and how policy gradient turns a sequence-level score into token-level updates.

From SFT to RL: The Two Degrees of Freedom

Why SFT is reference-distribution fitting, and how changing token weights and sampling turns it into RL.

Post-Training

From SFT to RL: Reward and Policy Gradient

Where RLHF's reward signal comes from, and how policy gradient turns a sequence-level score into token-level updates.

From SFT to RL: The Two Degrees of Freedom

Why SFT is reference-distribution fitting, and how changing token weights and sampling turns it into RL.