The Blog
Huayu’s Blog
Notes on reinforcement learning, AI for physics, and LLM post-training — written to clarify my own thinking, shared in case they help yours.
2026 2 posts
-
OPD 中 reverse KL 的 in Reward 与 in Loss
On-Policy Distillation 里的 reverse KL 恰好落在六宫格中"永远对"的 k1-in-Reward 格子;本文接着分析换成 k3、discount 设为 0、以及教师只暴露 top-k logprobs 这几个工程现实分别会踩什么坑。
-
LLM 强化学习中 KL 散度的正确形式是什么
从"无偏"和"有偏"的判断标准出发,拆解 k1/k2/k3 三种 KL 估计器在 KL-in-Reward 和 KL-in-Loss 两种实现下的梯度行为,说明"有偏性是估计器和放置位置组合的属性,不是估计器本身的属性"。