arXiv 2601.21968 · cs.CL

OVD: On-policy Verbal Distillation

Distill from a teacher you can only talk to. The teacher gives each step of the student's own rollout a score from 0 to 9. OVD keeps the steps that score well and asks the teacher to rewrite the rest. No teacher logits, no dense vocabulary tensors, and the student still explores its own trajectories.

Jing Xiong1, Hui Shen1, Shansan Gong1, Yuxin Cheng1, Jianghan Shen2, Chaofan Tao3, Haochen Tan3, Haoli Bai3, Lifeng Shang3, Ngai Wong1
1The University of Hong Kong   2Nanjing University   3Huawei Technologies
Paper arXiv Code (coming soon) BibTeX
Web Q&AWho was the British monarch when the host city of the 2012 Summer Olympics first hosted the Games? Gold answer: Edward VII. Search budget: 4 calls.
5
Student tokens (mS=1) Teacher-rewritten query (mT=1) Search result (masked from the loss)
Queries kept: – Search calls used: – Queries rewritten by teacher: – Outcome reward F1(y, y*): –
Figure 1. Verbal rejection sampling on one student rollout. Drag the threshold. In the default Web Q&A setting only the <search> queries are distilled. The teacher scores each query on a 0–9 scale, and a query scoring below τtr is replaced by one the teacher writes. The <think> steps are never scored or rewritten, so the reasoning stays the student's own. At low τtr the student keeps its vague queries, spends the search budget on irrelevant results, and guesses wrong. At τtr≥8 every query is the teacher's, so the student's own good queries never reach training. The middle range, including the paper's τtr=5, keeps good student queries and repairs only the weak ones. This rollout and its scores are made up to show how the method works. They do not come from the experiments.

TL;DR. Token-level on-policy distillation needs the teacher's full next-token distribution at every step. That is about 240 GB of logits for 32 rollouts at 8K tokens on a Qwen2.5-7B-class vocabulary, and closed-source teachers don't expose those numbers at all. OVD replaces logit matching with verbal ranking. A black-box teacher scores sub-trajectories, low-scoring ones are resampled by the teacher, and the resulting mixed trajectories train the student with GRPO. On eight Web Q&A benchmarks, OVD reaches 41.09% average exact match, +5.89 points over the strongest baseline.

Why not match the teacher's logits?

Standard on-policy distillation (OPD) samples trajectories from the student $\pi_S$ and minimizes the reverse KL to the teacher $\pi_{\mathcal E}$, one token at a time:

$$\mathrm{KL}(\pi_S\,\|\,\pi_{\mathcal E}) = \mathbb{E}\Big[\log \pi_S(x_\ell \mid q, x_{<\ell}) - \log \pi_{\mathcal E}(x_\ell \mid q, x_{<\ell})\Big].$$

This works well when you own the teacher. For long-horizon, agentic reasoning it breaks in two ways.

Storing dense student and teacher distributions over all $V$ tokens at every decoding step costs $\mathcal M = 1.5\cdot B\cdot N\cdot L\cdot V\cdot d_{32}$ bytes, since most stacks keep both FP32 and BF16 copies. That grows linearly in context length and in rollout count. Top-$k$ truncation cuts this by roughly $V/k$, but it still needs white-box access to the teacher.

8K
32
Figure 2. Logit storage for one problem (B=1). The calculation assumes a Qwen2.5-7B-class vocabulary (V≈152K), top-k with k=8,192, and a BF16 KV cache for reference (28 layers, 4 KV heads, head dim 128). The default, L=8K and N=32, matches the paper's Figure 1(b): about 240 GB for full-V logits against 13 GB for top-k. OVD stores no teacher logits. The teacher returns one score token per scored step.

Distilling search agents with top-1 and top-5 token-level OPD gives the pattern below. Exact match on the simulator's validation set keeps rising during training, but exact match with live search on held-out data stays far lower. Top-5 does worse than top-1 on both datasets. A denser token-level target seems to narrow the policy instead of guiding it.

PopQA

HotpotQA

Figure 3. Validation EM rises while held-out EM stays low. Solid lines show simulator validation EM over training. Dashed lines show held-out EM with live search. Top-5 matching (teal) ends below top-1 (orange) on held-out data.

Can a teacher's verbal score stand in for its logits?

OVD uses a verbal scoring interface. Given a problem $q$ and a student sub-trajectory $y$, the teacher outputs a single score token $S(y,q)\in\{0,1,\dots,9\}$. By default the score is decoded greedily, so a black-box API works. If the ten score-token logits happen to be available, the score can be sampled from them instead.

The first check is whether these scores carry signal. The teacher scores 32 fixed student prefixes for each of 128 prompts. The student then continues only from prefixes that score at least τte, and the final answers are graded.

Figure 4. Verbal scores pick better prefixes. This is final-answer EM from continuing the selected prefixes, compared with two baselines that select the same number of prefixes at random. The per-prompt baseline samples at random within each prompt, and the pooled baseline samples across all prompts. At τte=8, 441 prefixes remain. Verbal selection reaches 61.22%, against 49.06% for per-prompt random and 17.36% for pooled random. Shading shows 95% confidence intervals.

Verbal selection beats random selection even within a single prompt, so the score separates good prefixes from bad ones for the same question, beyond just flagging which questions are easy. OVD turns this signal into a sampler.

Verbal rejection sampling

OVD pipeline: the student produces on-policy rollouts; the teacher scores each step, accepts high-scoring ones, and resamples low-scoring ones; the resulting mixed trajectories are grouped, rewarded, and optimized with GRPO.
Figure 5. The OVD pipeline. Left: the student produces on-policy rollouts. Right: the teacher scores each step, accepts or resamples it, and the student resumes after each rewrite. The resulting mixed trajectories are grouped, rewarded against the reference answer, and optimized with GRPO.

Segmenting a rollout. In Web Q&A, the <think>, <search> and <information> tags mark natural step boundaries. Math rollouts have no tags. There OVD splits at the student's maximum-entropy token, a position that often marks a branch point in the reasoning.

Acceptance. For a threshold $\tau$, a student proposal is accepted with probability $a_\tau(y\mid q)=\Pr[S(y,q)\ge\tau]$. With greedy scoring this is simply $\mathbf 1[S\ge\tau]$. Conditioning on acceptance reweights the student's own distribution:

$$p_{a,\tau}(y\mid q)=\frac{\pi_S(y\mid q)\,a_\tau(y\mid q)}{Z_\tau(q)},\qquad Z_\tau(q)=\mathbb{E}_{y'\sim\pi_S}\big[a_\tau(y'\mid q)\big].$$

Replacement. A rejected step is resampled by the teacher. At a prefix $y_{<t}$, the next step therefore comes from a two-part mixture:

$$p_{m,\tau}(s_t\mid y_{

Takeaway. Thresholding turns verbal ranking into acceptance. The mixture approaches a teacher-preferred target $p_T$ when acceptance tracks the density ratio $p_T/\pi_S$ and the teacher's replacements are close to the target. The appendix makes this precise with total-variation bounds: $\mathrm{TV}(p_{m,\tau},p_T)\le \min\{1,\ \varepsilon+(1-Z_\tau)\eta\}$ per step, where $\varepsilon$ is the acceptance calibration error and $\eta$ the replacement error. Over a horizon of $T$ steps the per-step errors add up to at most $\sum_t\delta_t$.

Does acceptance actually track a density ratio? The paper checks against the teacher-to-student log-likelihood ratio, using $\pi_{\mathcal E}$ as a proxy for $p_T$, on 6,400 Math500 answers.

Figure 6. Acceptance rises with the teacher-to-student log-density ratio, then levels off. Points are within-question-centered bin estimates, and bars are 95% question-bootstrap intervals. Raising τ from 5 to 7 lowers acceptance substantially. The pattern favors answers the teacher finds relatively likely. On its own it doesn't prove the exact proportionality condition.

Training on mixed trajectories

Each complete trajectory receives an outcome reward: token-F1 against the reference for Web Q&A, exact match for math. Advantages are group-relative, as in GRPO. Every token records its source with $m^S_\ell$ or $m^T_\ell$ (where $m^S_\ell+m^T_\ell=1$), so the clipped objective splits cleanly:

$$\mathcal L_{\text{RL}}=\underbrace{-\mathbb{E}_m\big[m^S_\ell\, g_\ell(y)\big]}_{\text{learning from student exploration}}\;+\;\underbrace{-\mathbb{E}_m\big[m^T_\ell\, g_\ell(y)\big]}_{\text{learning from teacher guidance}},\qquad g_\ell=\min\!\big(\rho_\ell A,\ \mathrm{clip}(\rho_\ell,1\!-\!\epsilon_c,1\!+\!\epsilon_c)A\big).$$

A trajectory whose steps were all accepted contributes only the first term. The student learns from its own exploration without matching any teacher distribution. Search results are masked out of the loss.

Results: agentic Web Q&A

The student is Qwen-2.5-3B-Base and the teacher is Qwen2.5-7B, evaluated on eight benchmarks with live search at test time. Rows labeled OVD (τtr, τte) also use the teacher at inference with threshold τte. With τte=0, the student runs alone.

Table 1. Web Q&A exact match (%). Bold marks the best score in each column. ‡ marks process-reward-model baselines. * marks token-level OPD with full-vocabulary reverse KL or its top-k approximation. † samples the test-time score from the teacher's ten-token score distribution (temperature 0.7) instead of taking the most likely score.

With matched thresholds (5, 5), sampling the score instead of decoding it greedily raises average EM from 39.04% to 41.09%. Sampling implements the probabilistic acceptance rule $a_\tau=\Pr[S\ge\tau]$, and a greedy score reduces it to a hard gate. The two token-level OPD baselines reach only 14.14% and 11.51% average EM.

What should the teacher rewrite, and how often?

Table 2. Where to apply rejection. Macro-average EM with real search, (τtr, τte)=(5, 0), no teacher at test time.
Table 3. Test-time threshold. Average EM with τtr=5 fixed. A higher τte sends more steps to the teacher at inference. Single-hop and multi-hop columns are macro-averages over the per-dataset results.

Rejecting only search queries works best on single-hop questions. Rejecting only reasoning (CoT) steps works best on multi-hop questions. A moderate training threshold matters too. Replacing every student step (τtr=10) pushes the training mixture almost entirely toward teacher trajectories and discards student paths that might have been correct.

Web Q&A exact match per dataset under different training and testing rejection thresholds.
Figure 7. Training and test thresholds. Per-dataset and average EM for prompt-based teachers (dashed) and SFT-based teachers (solid) across (τtr, τte). In this comparison the paper finds that a moderate τte=5 outperforms τte=0 (accept everything) and τte=10 (teacher intervenes everywhere). In the appendix setting of Table 3 above, τte=10 has the highest average instead.

Results: tool-free math reasoning

To separate verbal distillation from tool use, the student is Qwen2.5-Math-1.5B and the teacher is DeepSeek-R1-Distill-Qwen-7B. Three ways of repairing a low-scoring rollout are compared:

  • OVD-FR (full response) resamples the whole trajectory.
  • OVD-HES (high-entropy suffix) keeps the prefix up to the student's max-entropy token and resamples only the suffix.
  • OVD-RS (random suffix) is a control that splits at a random position.
Figure 8. Macro-average accuracy on eight math benchmarks across checkpoints (AIME24, CARP, Middle, College, Gaokao-II, MATH, MAWPS, SAT-Math). The untrained student scores 33.15%. With 128 problems, OVD-HES has the best checkpoint overall (58.32% at step 300) but declines afterward. With one problem, it peaks at step 800 (56.26%). The paper stresses that these benchmarks were selected post hoc from 23 tasks, so the gaps are descriptive. OVD-RS results are reported only for the 128-problem setting.

Solution coverage (Pass@16)

Figure 9. Average P@16 on ASDiv, MAWPS, TABMWP, Olympiad and AIME24 (128 problems). At step 800, OVD-HES holds 78.50%, against 76.14% for OVD-RS and 75.26% for OVD-FR.

Time per training step (s)

Figure 10. Mean seconds per step against the number of training problems, on four AMD GPUs. With 128 problems, OVD-HES is 10.2% faster per step than OVD-FR.

RL tends to concentrate probability on a narrow set of reasoning paths. Keeping the student's own prefix and repairing only the weak continuation preserves late-stage solution coverage: +3.24 P@16 over full-response resampling at step 800. With 128 training problems it is also 10.2% faster per step than OVD-FR. With a single fixed problem, OVD-FR is the fastest.

Model checkpoints

Download the released checkpoints from Hugging Face. The math runs use Qwen2.5-Math-1.5B; the Web Q&A run uses Qwen2.5-3B-Base. Historical checkpoints are labeled as such in their model cards.

Browse all OVD checkpoints on Hugging Face

Math · one fixed problem

MethodStep 300Step 500Step 600Step 800
RLVR / GRPOModelModelModelModel
OVD-FR · Full responseModelModelModelModel
OVD-HES · High-entropy suffixModelModelModelModel
OVD-RS · Random suffixModelModelModelModel

Math · 128 distinct problems

MethodStep 300Step 500Step 600Step 800
RLVR / GRPOModelModelModelModel
OVD-FR · Full responseModelModelModelModel
OVD-HES · High-entropy suffixModelModelModelModel
OVD-RS · Random suffixModelModelModelModel

The 1-data RLVR step-300 checkpoint is a separate retraining run. The 128-data OVD-FR step-500 checkpoint is a model-only continuation; consult its model card for provenance.

Web Q&A

Base modelSettingCheckpointModel
Qwen2.5-3B-BasePure Query Reject · training threshold 5; think distillation offStep 200Download

The whole teacher interface

The teacher needs no fine-tuning and doesn't need to expose any probabilities. It answers one prompt with one number.

Verbal scoring prompt
Problem: q
Previous Steps: s₁, s₂, …, st−1
Current Step: st
Rate the quality of the current reasoning step on a scale of 0–9, where:
0–2Incorrect or irrelevant 3–5Partially correct but incomplete 6–8Correct and well-reasoned 9Optimal step toward the solution
Output only a single number from 0 to 9.

Contributions

  • Black-box on-policy distillation. OVD steers student rollouts using verbal scores and selective teacher continuations, without dense logits.
  • Distribution-approximation theory. Finite-horizon total-variation bounds on how close the mixed training distribution is to a teacher-preferred target, under explicit calibration and replacement-quality assumptions.
  • Evaluation across reasoning settings. Eight Web Q&A benchmarks and tool-free mathematics, with the checkpoint-dependent and resampling-dependent trade-offs reported.

BibTeX

@article{xiong2026ovd,
  title   = {{OVD}: On-policy Verbal Distillation},
  author  = {Xiong, Jing and Shen, Hui and Gong, Shansan and Cheng, Yuxin and
             Shen, Jianghan and Tao, Chaofan and Tan, Haochen and Bai, Haoli and
             Shang, Lifeng and Wong, Ngai},
  journal = {arXiv preprint arXiv:2601.21968},
  year    = {2026}
}