PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

Youzhi Liu, Ruobing Zheng*, Boyuan Tong, Tianqi Li,
Pingqi Li, Hanbo Bi, Yi Yuan, Jingdong Chen
Ant Group
*Corresponding author

Task-preferred update directions and optimization trajectories of MOPD and PMOPD.

Task-update geometry. Math, Code, and Reason prefer different update directions at the same checkpoint. PMOPD corrects updates against protected task directions and follows a more stable trajectory toward a shared low-loss region.

Abstract

Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single language model, but updates from different teachers can interfere in shared parameters. We find that cumulative updates from different tasks rapidly concentrate in distinct low-dimensional subspaces. Based on this geometry, we propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from realized task-block displacements and projects both gradients and optimizer updates to remove components that interfere with protected task directions. A lightweight conflict probe guides task ordering, while a cycling strategy balances reliable subspace estimation with timely task revisitation. Across Code, Reason, and Math, PMOPD improves every evaluated capability over MOPD, increasing the three-task average by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B.

Method

Subspace memory construction and projected protection in multi-teacher OPD.

Overview of the PMOPD method across Code, Reason, and Math stages.

PMOPD overview. Within each cycle, completed task blocks are summarized by the dominant singular vectors of their cumulative parameter displacements. The resulting orthogonal memory constrains subsequent task gradients and preconditioned optimizer updates. Memory is rebuilt in every cycle.

PMOPD trains the shared student in ordered task blocks. After a task block, it computes the realized parameter displacement between the block's start and end checkpoints and extracts its top singular directions. These directions are orthogonalized into a compact protected memory. During later task blocks, PMOPD removes update components that lie in this memory. The projection is applied twice: first to the raw gradient and then to the preconditioned Adafactor update, because adaptive element-wise scaling can rotate a projected gradient back toward a protected subspace.

Task Ordering and Cycling

Task interference is directional, and the final result depends on both the ordering of teachers and how frequently the model revisits each task. PMOPD uses lightweight geometric diagnostics instead of exhaustive full-training searches.

Ten-fold estimates of task-level undirected conflict for Code, Reason, and Math.

Conflict probe. Ten estimates based on 60-sample subsets consistently recover the ordering Code < Reason < Math, selecting Code→Reason→Math for training.

Average score and cross-cycle subspace consistency for different cycle counts.

Cycle selection. Both the average task score and cross-cycle subspace consistency peak at four cycles, balancing reliable block-displacement estimates with timely task revisitation.

Main Results

PMOPD improves Math, Reason, and Code simultaneously on two independently developed model families. The gains therefore reflect stronger joint capability integration rather than a transfer of performance from one task to another.

Method Qwen2.5-7B Llama-3.1-8B
MathReasonCodeAvg. MathReasonCodeAvg.
Student Model51.7850.2049.6750.5511.3358.5627.3332.41
Parameter Merge61.5660.4453.6758.5612.2263.6430.0035.29
MOPD65.5671.7456.0064.4318.2268.6330.0038.95
BB-MOPD65.3370.6855.6763.8915.3364.2129.3336.29
Open-MOPD66.4470.5253.6763.5414.8965.5225.0035.14
PMOPD67.7873.4659.6766.9721.3369.4532.3341.04

PMOPD improves the average over MOPD by +2.54 points on Qwen2.5-7B and +2.09 points on Llama-3.1-8B.

BibTeX

@article{liu2026pmopd,
  title={PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation},
  author={Liu, Youzhi and Zheng, Ruobing and Tong, Boyuan and Li, Tianqi and Li, Pingqi and Bi, Hanbo and Yuan, Yi and Chen, Jingdong},
  journal={arXiv preprint arXiv:2609.34605},
  year={2026}
}