OLIVE Distillation Method Improves Reasoning Performance via Teacher Continuations
September 30, 2026
OLIVE uses teacher-generated continuations on student-produced prefixes to address covariate shift and prefix failure in on-policy distillation. This method achieves higher reasoning performance than token-level on-policy distillation (OPD) with a 23.8% reduction in total training time via asynchronous implementation.
HOW THIS AFFECTS YOU
●
builderYou can reduce training time and improve reasoning capabilities during distillation using this asynchronous approach.
●
researcherThis provides a way to bypass the limitations of fixed teacher trajectories and token-level prefix failures.