Blog

Research / October 2, 2026

Sharpening the spear

Post-training sharpens a model: more right on the first try, fewer problems it can solve at all. A new paper calls that loss the Sharpening Tax and proposes a fix: sample hard problems hotter during training. We tested it on math against plain RL and a method that practices more. The fix did not help. More practice did.

The short version. Think of training a model as making a spear. Pre-training forges the metal. Mid-training adds more. Post-training sharpens the point. Sharpening removes metal: the model gets more answers right on the first try and can reach fewer answers at all. A new paper measures that loss and proposes a fix that turns the model's randomness up on problems it keeps failing. We tested the fix on competition math. It matched plain training. A method that simply practices hard problems more often did better, at almost four times the compute.

Math problems solved on the first try, with 95% intervals

320 held-out MATH-500 problems at levels 3 to 5, 8 attempts each, Qwen2.5-1.5B. Bars are the first training seed; re-running the untrained eval three times moved it by under a point.

Math problems solved on the first try, with 95% intervals
ItemValue95% interval
Untrained model20.2%17.3% to 23.2%
Plain RL (GRPO)30.8%26.9% to 34.7%
Temperature fix (PTGS)31.2%27.3% to 35.1%
Practice more (Reinforce-Ada)33.9%29.9% to 38.1%

Sharpening removes metal

Reinforcement learning (RL) on a checker trains a model on its own attempts. It pushes up the attempts that passed and pushes down the ones that failed [3]. Over time the model stops wandering. Problems it half-knew become problems it always solves. That is the point of training, and it has a cost the paper names: some half-known problems go the other way and become problems it never solves [1]. The share it can solve in many tries (pass@8 here, the "coverage") stops growing or shrinks. Earlier work saw this in math and code [4]. The new paper sees it in agents too, and finds that larger models pay more.

The Sharpening Tax is one number for that loss. It counts how much a model gains from extra tries, before and after training. A positive tax means training took away the value of trying again.

Two ways to keep the metal

Both fixes target the same waste. RL learns by comparing attempts at one problem. If all four attempts fail, or all four pass, there is nothing to compare, and that problem teaches nothing that step.

  • Turn the heat up (PTGS) [1]. Keep a running guess of how often the model solves each problem. Sample the hard ones at a higher temperature, so the model tries less likely answers, and the easy ones at a lower one. Same number of attempts.
  • Practice more (Reinforce-Ada) [2]. Keep sampling a problem until it has two passes and two failures, up to 32 tries, then train on four of them.
Training steps where a problem taught nothing

Share of problems per training step whose four trained attempts all passed or all failed, mean over 80 steps and both seeds.

Training steps where a problem taught nothing
ItemValue
Plain RL (GRPO)68%
Temperature fix (PTGS)71%
Practice more (Reinforce-Ada)41%

What happened

The temperature fix did not reduce the wasted problems. It added a few. Heat only helps if a problem the model fails at normal temperature starts passing at a higher one, and on these math problems it mostly did not. Cooling the easy problems made them pass every time, which teaches nothing either. On first-try accuracy it tied plain RL on one seed and beat it by four points on the other. The difference is not proven.

Practicing more cut the wasted problems almost in half. Both of its runs beat both plain RL runs on first try and on eight tries. It also took 3.7 times as long to train.

No method paid a Sharpening Tax. Every trained model solved more problems in eight tries than the untrained one. The paper finds the tax on fully post-trained models with up to 128 tries. Eighty small training steps on a 1.5B model did not sharpen it that far.

What this teaches about post-training

Hard problems need more attempts, not hotter ones. The spear only loses metal when you grind it long enough. To measure the loss, keep the untrained model as a reference and score both at many tries.

For researchers

SettingValue
ModelQwen/Qwen2.5-1.5B-Instruct, LoRA r=32, TRL 0.19.1 GRPO, lr 1e-4, no KL, on-policy
DataMATH train levels 3 to 5, 192 prompts (5 visits each); MATH-500 levels 3 to 5, 320 held out
Update80 steps x 12 prompts x 4 rollouts in every arm; reward = Math-Verify against the answer
PTGStau 1.5, forgetting 0.95, target success 0.25 to 0.5, prior mass 2, log-probs at each row's temperature [1]
Reinforce-Adabalanced exit, rounds of 8, at most 32 draws, keep 4, pool-rate baseline [2]
pass@1 vs GRPOPTGS +0.4 [-1.7, +2.5]; Reinforce-Ada +3.1 [+0.7, +5.4] points (seed-17 pair, 320 tasks)
pass@8 vs GRPOPTGS +3.9 [+1.1, +6.9]; Reinforce-Ada +7.2 [+3.6, +10.6] points (task bootstrap, seeds averaged)
Tax_S(8) vs baseGRPO +0.006, PTGS -0.007, Reinforce-Ada -0.013; every interval covers zero
Noise floorbase re-run three times, run_std 0.8 points; a single-pair delta under 5.1 points reads flat
Cost804 GPU minutes on L40S, $26.79; training 51 to 57 minutes a seed, Reinforce-Ada 185 to 207

Run it

pip install whileai
git clone https://github.com/whilehq/whileai-sdk
cd whileai-sdk/recipes/papers/ptgs
python recipe.py --selftest      # the temperature rule and the tax, no GPU
python recipe.py                 # three arms, two seeds, on Modal

References

  1. Oh, C., Zeng, Q., Qi, Q., Zhmoginov, A., Lei, D., He, Y., Phan, H., Kang, H., Mirhoseini, A. and Li, S. (2026). Sharpening tax in post-training. arXiv:2610.01509.
  2. Xiong, W., et al. (2025). Reinforce-Ada: An adaptive sampling framework under non-linear RL objectives. arXiv:2510.04996.
  3. Shao, Z., et al. (2024). DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv:2402.03300.
  4. Yue, Y., et al. (2025). Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? arXiv:2504.13837.