Representation Learning
Studying what makes a visual representation universal, useful, and "good".
Studying what makes a visual representation universal, useful, and "good".
Hardware-software co-design across speculative decoding, quantization, and pruning.
Building agent tooling and constructing benchmarks for evaluating agent capabilities.
Where RLHF's reward signal comes from, and how policy gradient turns a sequence-level score into token-level updates.
Why SFT is reference-distribution fitting, and how changing token weights and sampling turns it into RL.