PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
Youzhi Liu,
Ruobing Zheng*,
Boyuan Tong,
Tianqi Li,
Pingqi Li,
Hanbo Bi,
Yi Yuan,
Jingdong Chen
Ant Group
*Corresponding author
Abstract
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single language model, but updates from different teachers can interfere in shared parameters. We find that cumulative updates from different tasks rapidly concentrate in distinct low-dimensional subspaces. Based on this geometry, we propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from realized task-block displacements and projects both gradients and optimizer updates to remove components that interfere with protected task directions. A lightweight conflict probe guides task ordering, while a cycling strategy balances reliable subspace estimation with timely task revisitation. Across Code, Reason, and Math, PMOPD improves every evaluated capability over MOPD, increasing the three-task average by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B.
Method
Subspace memory construction and projected protection in multi-teacher OPD.
PMOPD trains the shared student in ordered task blocks. After a task block, it computes the realized parameter displacement between the block's start and end checkpoints and extracts its top singular directions. These directions are orthogonalized into a compact protected memory. During later task blocks, PMOPD removes update components that lie in this memory. The projection is applied twice: first to the raw gradient and then to the preconditioned Adafactor update, because adaptive element-wise scaling can rotate a projected gradient back toward a protected subspace.
Task Ordering and Cycling
Task interference is directional, and the final result depends on both the ordering of teachers and how frequently the model revisits each task. PMOPD uses lightweight geometric diagnostics instead of exhaustive full-training searches.
Main Results
PMOPD improves Math, Reason, and Code simultaneously on two independently developed model families. The gains therefore reflect stronger joint capability integration rather than a transfer of performance from one task to another.
| Method |
Qwen2.5-7B |
Llama-3.1-8B |
| Math | Reason | Code | Avg. |
Math | Reason | Code | Avg. |
| Student Model | 51.78 | 50.20 | 49.67 | 50.55 | 11.33 | 58.56 | 27.33 | 32.41 |
| Parameter Merge | 61.56 | 60.44 | 53.67 | 58.56 | 12.22 | 63.64 | 30.00 | 35.29 |
| MOPD | 65.56 | 71.74 | 56.00 | 64.43 | 18.22 | 68.63 | 30.00 | 38.95 |
| BB-MOPD | 65.33 | 70.68 | 55.67 | 63.89 | 15.33 | 64.21 | 29.33 | 36.29 |
| Open-MOPD | 66.44 | 70.52 | 53.67 | 63.54 | 14.89 | 65.52 | 25.00 | 35.14 |
| PMOPD | 67.78 | 73.46 | 59.67 | 66.97 | 21.33 | 69.45 | 32.33 | 41.04 |
PMOPD improves the average over MOPD by +2.54 points on Qwen2.5-7B and +2.09 points on Llama-3.1-8B.
BibTeX
@article{liu2026pmopd,
title={PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation},
author={Liu, Youzhi and Zheng, Ruobing and Tong, Boyuan and Li, Tianqi and Li, Pingqi and Bi, Hanbo and Yuan, Yi and Chen, Jingdong},
journal={arXiv preprint arXiv:2609.34605},
year={2026}
}