Research / October 2, 2026
Sharpening the spear
Post-training sharpens a model: more right on the first try, fewer problems it can solve at all. A new paper calls that loss the Sharpening Tax and proposes a fix: sample hard problems hotter during training. We tested it on math against plain RL and a method that practices more. The fix did not help. More practice did.
The short version. Think of training a model as making a spear. Pre-training forges the metal. Mid-training adds more. Post-training sharpens the point. Sharpening removes metal: the model gets more answers right on the first try and can reach fewer answers at all. A new paper measures that loss and proposes a fix that turns the model's randomness up on problems it keeps failing. We tested the fix on competition math. It matched plain training. A method that simply practices hard problems more often did better, at almost four times the compute.
320 held-out MATH-500 problems at levels 3 to 5, 8 attempts each, Qwen2.5-1.5B. Bars are the first training seed; re-running the untrained eval three times moved it by under a point.
| Item | Value | 95% interval |
|---|---|---|
| Untrained model | 20.2% | 17.3% to 23.2% |
| Plain RL (GRPO) | 30.8% | 26.9% to 34.7% |
| Temperature fix (PTGS) | 31.2% | 27.3% to 35.1% |
| Practice more (Reinforce-Ada) | 33.9% | 29.9% to 38.1% |
Sharpening removes metal
Reinforcement learning (RL) on a checker trains a model on its own attempts. It pushes up the attempts that passed and pushes down the ones that failed [3]. Over time the model stops wandering. Problems it half-knew become problems it always solves. That is the point of training, and it has a cost the paper names: some half-known problems go the other way and become problems it never solves [1]. The share it can solve in many tries (pass@8 here, the "coverage") stops growing or shrinks. Earlier work saw this in math and code [4]. The new paper sees it in agents too, and finds that larger models pay more.
The Sharpening Tax is one number for that loss. It counts how much a model gains from extra tries, before and after training. A positive tax means training took away the value of trying again.
Two ways to keep the metal
Both fixes target the same waste. RL learns by comparing attempts at one problem. If all four attempts fail, or all four pass, there is nothing to compare, and that problem teaches nothing that step.
- Turn the heat up (PTGS) [1]. Keep a running guess of how often the model solves each problem. Sample the hard ones at a higher temperature, so the model tries less likely answers, and the easy ones at a lower one. Same number of attempts.
- Practice more (Reinforce-Ada) [2]. Keep sampling a problem until it has two passes and two failures, up to 32 tries, then train on four of them.
Share of problems per training step whose four trained attempts all passed or all failed, mean over 80 steps and both seeds.
| Item | Value |
|---|---|
| Plain RL (GRPO) | 68% |
| Temperature fix (PTGS) | 71% |
| Practice more (Reinforce-Ada) | 41% |
What happened
The temperature fix did not reduce the wasted problems. It added a few. Heat only helps if a problem the model fails at normal temperature starts passing at a higher one, and on these math problems it mostly did not. Cooling the easy problems made them pass every time, which teaches nothing either. On first-try accuracy it tied plain RL on one seed and beat it by four points on the other. The difference is not proven.
Practicing more cut the wasted problems almost in half. Both of its runs beat both plain RL runs on first try and on eight tries. It also took 3.7 times as long to train.
No method paid a Sharpening Tax. Every trained model solved more problems in eight tries than the untrained one. The paper finds the tax on fully post-trained models with up to 128 tries. Eighty small training steps on a 1.5B model did not sharpen it that far.
What this teaches about post-training
Hard problems need more attempts, not hotter ones. The spear only loses metal when you grind it long enough. To measure the loss, keep the untrained model as a reference and score both at many tries.
For researchers
| Setting | Value |
|---|---|
| Model | Qwen/Qwen2.5-1.5B-Instruct, LoRA r=32, TRL 0.19.1 GRPO, lr 1e-4, no KL, on-policy |
| Data | MATH train levels 3 to 5, 192 prompts (5 visits each); MATH-500 levels 3 to 5, 320 held out |
| Update | 80 steps x 12 prompts x 4 rollouts in every arm; reward = Math-Verify against the answer |
| PTGS | tau 1.5, forgetting 0.95, target success 0.25 to 0.5, prior mass 2, log-probs at each row's temperature [1] |
| Reinforce-Ada | balanced exit, rounds of 8, at most 32 draws, keep 4, pool-rate baseline [2] |
| pass@1 vs GRPO | PTGS +0.4 [-1.7, +2.5]; Reinforce-Ada +3.1 [+0.7, +5.4] points (seed-17 pair, 320 tasks) |
| pass@8 vs GRPO | PTGS +3.9 [+1.1, +6.9]; Reinforce-Ada +7.2 [+3.6, +10.6] points (task bootstrap, seeds averaged) |
| Tax_S(8) vs base | GRPO +0.006, PTGS -0.007, Reinforce-Ada -0.013; every interval covers zero |
| Noise floor | base re-run three times, run_std 0.8 points; a single-pair delta under 5.1 points reads flat |
| Cost | 804 GPU minutes on L40S, $26.79; training 51 to 57 minutes a seed, Reinforce-Ada 185 to 207 |
Run it
pip install whileai
git clone https://github.com/whilehq/whileai-sdk
cd whileai-sdk/recipes/papers/ptgs
python recipe.py --selftest # the temperature rule and the tax, no GPU
python recipe.py # three arms, two seeds, on ModalReferences
- Oh, C., Zeng, Q., Qi, Q., Zhmoginov, A., Lei, D., He, Y., Phan, H., Kang, H., Mirhoseini, A. and Li, S. (2026). Sharpening tax in post-training. arXiv:2610.01509.
- Xiong, W., et al. (2025). Reinforce-Ada: An adaptive sampling framework under non-linear RL objectives. arXiv:2510.04996.
- Shao, Z., et al. (2024). DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv:2402.03300.
- Yue, Y., et al. (2025). Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? arXiv:2504.13837.