Research / September 28, 2026
Deep Agents SFT from traces
We fine-tuned a LangChain Deep Agents code reviewer twice on its own LangSmith traces, with the same smithtune pipeline. Only the choice of traces differed. Trained on the correct, short traces While. chose, it got 57% of SWE-bench Verified reviews right against 46%, with fewer model calls. Each arm has one training run so far.
The short version. We wanted to know if the choice of training examples matters as much as the trainer. A LangChain Deep Agents agent reviewed code patches and its runs went to LangSmith. We fine-tuned it twice with LangChain's smithtune, changing only which runs it learned from. Trained on correct, short runs, it got 57% of reviews right on a public benchmark, against 46% when trained on every correct run. It also took fewer steps. On unseen repos the two tied. Each arm has one training run so far, and a second is next.
Every arm served from one vLLM server with thinking off and the same sampling, so only the weights differ. The three base runs set the noise floor at 7.8 points.
| Item | Value |
|---|---|
| Base model, run 1 | 11.6% |
| Base model, run 2 | 12.0% |
| Base model, run 3 | 9.6% |
| Tuned harness profile, no training | 20.4% |
| SFT on every correct trace | 46.0% |
| SFT on correct and lean traces | 56.8% |
Six steps from traces to a tuned reviewer
- Run the stock agent and keep its runs. A stock Deep Agents agent [1] on Qwen3.8-27B, an open model, reads a proposed patch for a real GitHub issue with read-only file tools and answers approve or reject. Each run is recorded step by step (a trace) in LangSmith [2] with a score, correct or not. Hidden tests already decided whether each patch fixed its issue, so the right answer is known. We collected 400 training reviews.
- Choose the training traces two ways. Without While., a panel of model judges (smithtune's judge council) [3] reads each trace against a plain rubric. It kept 258 of 260 correct traces. With While., a trace is kept only if it is correct and used at most eight model calls, spread across the different ways the agent used its tools. That kept 173.
- Fine-tune both the same way. The model learns to copy the kept examples (supervised fine-tuning, SFT), with smithtune's own renderer, split and schedule.
- Serve both from one server. Same sampling, thinking off, so any gap comes from the weights.
- Score on a benchmark nobody here built. SWE-bench Verified [4] patch review has 250 reviews on 125 issues, and half the patches fixed their issue. None of its repos appear in the training traces.
- Measure the noise first. A single base run is not a baseline. Re-running the unchanged base once moved it 4.5 points, with an interval that excluded zero. So every comparison here carries how far the score moves when nothing changes (the noise floor), taken from three base runs.
Lean traces taught a better reviewer
The base model mostly ran out of steps before deciding. Tuning the Deep Agents settings (its harness), with no training, lifted it by about nine points. The model trained on While.'s traces was right more often, with a gain between +5 and +17 points, the range the true gain very likely sits in. That clears the noise floor. It also used about five fewer model calls per review.
84 repositories the training traces never touched, same server and sampling. The two fine-tunes are a tie, so the difference could be chance.
| Item | Value |
|---|---|
| Base model, run 1 | 33.3% |
| Base model, run 2 | 27.2% |
| Base model, run 3 | 26.4% |
| Tuned harness profile, no training | 35.0% |
| SFT on every correct trace | 63.8% |
| SFT on correct and lean traces | 61.4% |
Mean model calls per review for the two fine-tunes. The model trained on lean traces is cheaper on both tests, and both gaps clear zero.
| Item | Value |
|---|---|
| SWE-bench Verified, every correct trace | 23.5 |
| SWE-bench Verified, correct and lean | 18.2 |
| Held-out repos, every correct trace | 14.2 |
| Held-out repos, correct and lean | 11.1 |
On the second test, repos the model never trained on (held-out repos), accuracy was not proven either way. The lean model still used about three fewer calls per review. Its training set was a third smaller and took 65% fewer training tokens.
What this teaches about post-training
Which traces you train on matters as much as the trainer. Fine-tuning copies behavior, all of it. Keeping every correct trace also teaches the agent the long, wandering ones. A correct review that took twenty calls is a lesson in taking twenty calls.
Filtering sampled outputs by a score and training on the survivors is rejection sampling [5]. Adding a budget on model calls turns efficiency into data. Choosing correct and lean traces taught a model that is more accurate on unfamiliar large repos such as django and sympy, and cheaper to run, from a third less data.
The pipeline stayed LangChain's. Deep Agents ran the agent, LangSmith kept the traces and smithtune trained the model. While. decided what to train on and whether it worked.
Three runs come next. A second training seed per arm. A random set the size of the lean set drawn from the council's set, to separate the choice of traces from their number. And the base with thinking on, which scored between 70 and 75% on the held-out repos, on SWE-bench Verified. The tuned models are non-thinking because smithtune leaves reasoning out of training rows by default.
For researchers
| Setting | Value |
|---|---|
| Agent | LangChain Deep Agents create_deep_agent, read-only filesystem over a tarball of the base commit, Qwen3.8-27B |
| Training pool | 400 reviews built from nebius/SWE-agent-trajectories and nebius/SWE-bench-extra [6], split by repository, 67% correct |
| Council traces | smithtune dataset pull with the correctness filter, then triage with rubric.md, 258 of 260 kept |
| Lean traces | wai.optimize(mode="sft"), reward 1 only for a correct verdict within 8 model calls, round-robin over tool-call signatures, 173 kept |
| Rows | smithtune prepare_sft_rows (reasoning omitted, its default), split_rows 80/10/10, render_row_tokens (one datum per assistant turn) |
| Train datums | 741 lean, 1,606 council |
| Tokens per epoch | 3,619,086 lean, 10,221,451 council |
| Schedule | LoRA rank 8, batch 32, lr 1e-4, seed 42, up to 5 epochs, early stop on validation loss, both kept epoch 2 |
| Hardware | one Modal H200 per arm |
| Serving | one vLLM 0.26 server, base and both adapters, enable_thinking=false, same sampling |
| Public benchmark | SWE-bench Verified [4], 250 reviews on 125 issues, 50/50 approve and reject, six public leaderboard submissions, none of its twelve repositories in training |
| Held-out repos | 246 reviews on 84 repositories, wai.decontaminate dropped 0 of 1,306 |
| Paired test | wai.compare_runs, paired by review, 95% bootstrap interval |
| Noise floor | wai.eval_variance over three base runs, t at 2 df, 7.8 points public, 23.0 held-out |
| Base, thinking on | 70.3, 75.1 and 74.6% on held-out repos via OpenRouter (third run 213 reviews) |
The metric is accuracy, per review, and the comparison is the paired mean difference with a bootstrap interval over reviews:
The public gain clears the 7.8-point noise floor. The held-out interval spans zero.
| Item | Value | 95% interval |
|---|---|---|
| SWE-bench Verified, n = 250 | +10.8 pts | +4.8 pts to +16.8 pts |
| Held-out repos, n = 246 | -2.4 pts | -7.3 pts to +2.4 pts |
Negative is fewer calls. Both intervals exclude zero.
| Item | Value | 95% interval |
|---|---|---|
| SWE-bench Verified, n = 250 | -5.3 | -6.7 to -4.0 |
| Held-out repos, n = 246 | -3.1 | -4.3 to -2.0 |
Every number is in the recipe's results.json [7].
References
- LangChain (2025). Deep Agents [software]. github.com/langchain-ai/deepagents. Accessed September 28, 2026.
- LangChain (2026). LangSmith documentation. docs.langchain.com/langsmith. Accessed September 28, 2026.
- LangChain (2026). smithtune v0.1.0 [software]. github.com/langchain-ai/smithtune. Accessed September 28, 2026.
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O. and Narasimhan, K. (2024). SWE-bench: Can language models resolve real-world GitHub issues? ICLR 2024. arXiv:2310.06770. SWE-bench Verified subset: OpenAI (2024), Introducing SWE-bench Verified, openai.com/index/introducing-swe-bench-verified.
- Lambert, N. (2025). Reinforcement Learning from Human Feedback, chapter Rejection Sampling. arXiv:2504.12501. Online at rlhfbook.com.
- Nebius (2024). SWE-agent-trajectories and SWE-bench-extra [datasets]. huggingface.co/datasets/nebius/SWE-agent-trajectories.
- whilehq (2026). deepagents-review-four-arms recipe, whileai SDK [code and data]. github.com/whilehq/whileai-sdk/tree/main/recipes/community/deepagents-review-four-arms.
Run it
pip install whileai
git clone https://github.com/whilehq/whileai-sdk
cd whileai-sdk/recipes/community/deepagents-review-four-arms
python run.py --dry-run # the comparison and the selection rule on fixture rows, no keys
python run.py traces # stock agent over the training reviews, traces to LangSmith
python pick.py # the lean traces, written back as LangSmith feedback
python make_results.py # paired comparisons and the noise floor