Blog

Research / September 28, 2026

Deep Agents SFT from traces

We fine-tuned a LangChain Deep Agents code reviewer twice on its own LangSmith traces, with the same smithtune pipeline. Only the choice of traces differed. Trained on the correct, short traces While. chose, it got 57% of SWE-bench Verified reviews right against 46%, with fewer model calls. Each arm has one training run so far.

The short version. We wanted to know if the choice of training examples matters as much as the trainer. A LangChain Deep Agents agent reviewed code patches and its runs went to LangSmith. We fine-tuned it twice with LangChain's smithtune, changing only which runs it learned from. Trained on correct, short runs, it got 57% of reviews right on a public benchmark, against 46% when trained on every correct run. It also took fewer steps. On unseen repos the two tied. Each arm has one training run so far, and a second is next.

SWE-bench Verified patch review, share of 250 reviews correct

Every arm served from one vLLM server with thinking off and the same sampling, so only the weights differ. The three base runs set the noise floor at 7.8 points.

SWE-bench Verified patch review, share of 250 reviews correct
ItemValue
Base model, run 111.6%
Base model, run 212.0%
Base model, run 39.6%
Tuned harness profile, no training20.4%
SFT on every correct trace46.0%
SFT on correct and lean traces56.8%

Six steps from traces to a tuned reviewer

  1. Run the stock agent and keep its runs. A stock Deep Agents agent [1] on Qwen3.8-27B, an open model, reads a proposed patch for a real GitHub issue with read-only file tools and answers approve or reject. Each run is recorded step by step (a trace) in LangSmith [2] with a score, correct or not. Hidden tests already decided whether each patch fixed its issue, so the right answer is known. We collected 400 training reviews.
  2. Choose the training traces two ways. Without While., a panel of model judges (smithtune's judge council) [3] reads each trace against a plain rubric. It kept 258 of 260 correct traces. With While., a trace is kept only if it is correct and used at most eight model calls, spread across the different ways the agent used its tools. That kept 173.
  3. Fine-tune both the same way. The model learns to copy the kept examples (supervised fine-tuning, SFT), with smithtune's own renderer, split and schedule.
  4. Serve both from one server. Same sampling, thinking off, so any gap comes from the weights.
  5. Score on a benchmark nobody here built. SWE-bench Verified [4] patch review has 250 reviews on 125 issues, and half the patches fixed their issue. None of its repos appear in the training traces.
  6. Measure the noise first. A single base run is not a baseline. Re-running the unchanged base once moved it 4.5 points, with an interval that excluded zero. So every comparison here carries how far the score moves when nothing changes (the noise floor), taken from three base runs.

Lean traces taught a better reviewer

The base model mostly ran out of steps before deciding. Tuning the Deep Agents settings (its harness), with no training, lifted it by about nine points. The model trained on While.'s traces was right more often, with a gain between +5 and +17 points, the range the true gain very likely sits in. That clears the noise floor. It also used about five fewer model calls per review.

Repos kept out of training, share of 246 reviews correct

84 repositories the training traces never touched, same server and sampling. The two fine-tunes are a tie, so the difference could be chance.

Repos kept out of training, share of 246 reviews correct
ItemValue
Base model, run 133.3%
Base model, run 227.2%
Base model, run 326.4%
Tuned harness profile, no training35.0%
SFT on every correct trace63.8%
SFT on correct and lean traces61.4%
Model calls per review, fewer is cheaper to run

Mean model calls per review for the two fine-tunes. The model trained on lean traces is cheaper on both tests, and both gaps clear zero.

Model calls per review, fewer is cheaper to run
ItemValue
SWE-bench Verified, every correct trace23.5
SWE-bench Verified, correct and lean18.2
Held-out repos, every correct trace14.2
Held-out repos, correct and lean11.1

On the second test, repos the model never trained on (held-out repos), accuracy was not proven either way. The lean model still used about three fewer calls per review. Its training set was a third smaller and took 65% fewer training tokens.

What this teaches about post-training

Which traces you train on matters as much as the trainer. Fine-tuning copies behavior, all of it. Keeping every correct trace also teaches the agent the long, wandering ones. A correct review that took twenty calls is a lesson in taking twenty calls.

Filtering sampled outputs by a score and training on the survivors is rejection sampling [5]. Adding a budget on model calls turns efficiency into data. Choosing correct and lean traces taught a model that is more accurate on unfamiliar large repos such as django and sympy, and cheaper to run, from a third less data.

The pipeline stayed LangChain's. Deep Agents ran the agent, LangSmith kept the traces and smithtune trained the model. While. decided what to train on and whether it worked.

Three runs come next. A second training seed per arm. A random set the size of the lean set drawn from the council's set, to separate the choice of traces from their number. And the base with thinking on, which scored between 70 and 75% on the held-out repos, on SWE-bench Verified. The tuned models are non-thinking because smithtune leaves reasoning out of training rows by default.

For researchers

SettingValue
AgentLangChain Deep Agents create_deep_agent, read-only filesystem over a tarball of the base commit, Qwen3.8-27B
Training pool400 reviews built from nebius/SWE-agent-trajectories and nebius/SWE-bench-extra [6], split by repository, 67% correct
Council tracessmithtune dataset pull with the correctness filter, then triage with rubric.md, 258 of 260 kept
Lean traceswai.optimize(mode="sft"), reward 1 only for a correct verdict within 8 model calls, round-robin over tool-call signatures, 173 kept
Rowssmithtune prepare_sft_rows (reasoning omitted, its default), split_rows 80/10/10, render_row_tokens (one datum per assistant turn)
Train datums741 lean, 1,606 council
Tokens per epoch3,619,086 lean, 10,221,451 council
ScheduleLoRA rank 8, batch 32, lr 1e-4, seed 42, up to 5 epochs, early stop on validation loss, both kept epoch 2
Hardwareone Modal H200 per arm
Servingone vLLM 0.26 server, base and both adapters, enable_thinking=false, same sampling
Public benchmarkSWE-bench Verified [4], 250 reviews on 125 issues, 50/50 approve and reject, six public leaderboard submissions, none of its twelve repositories in training
Held-out repos246 reviews on 84 repositories, wai.decontaminate dropped 0 of 1,306
Paired testwai.compare_runs, paired by review, 95% bootstrap interval
Noise floorwai.eval_variance over three base runs, t at 2 df, 7.8 points public, 23.0 held-out
Base, thinking on70.3, 75.1 and 74.6% on held-out repos via OpenRouter (third run 213 reviews)

The metric is accuracy, yi∈{0,1}y_i \in \{0, 1\} per review, and the comparison is the paired mean difference with a bootstrap interval over reviews:

Δ^=1n∑i=1n(yilean−yicouncil),floor=t0.975, 2 2 sbase\hat{\Delta} = \frac{1}{n}\sum_{i=1}^{n}\left(y_i^{\text{lean}} - y_i^{\text{council}}\right), \qquad \text{floor} = t_{0.975,\,2}\,\sqrt{2}\,s_{\text{base}}
Lean traces minus council traces, accuracy, paired by review

The public gain clears the 7.8-point noise floor. The held-out interval spans zero.

Lean traces minus council traces, accuracy, paired by review
ItemValue95% interval
SWE-bench Verified, n = 250+10.8 pts+4.8 pts to +16.8 pts
Held-out repos, n = 246-2.4 pts-7.3 pts to +2.4 pts
Lean traces minus council traces, model calls per review, paired

Negative is fewer calls. Both intervals exclude zero.

Lean traces minus council traces, model calls per review, paired
ItemValue95% interval
SWE-bench Verified, n = 250-5.3-6.7 to -4.0
Held-out repos, n = 246-3.1-4.3 to -2.0

Every number is in the recipe's results.json [7].

References

  1. LangChain (2025). Deep Agents [software]. github.com/langchain-ai/deepagents. Accessed September 28, 2026.
  2. LangChain (2026). LangSmith documentation. docs.langchain.com/langsmith. Accessed September 28, 2026.
  3. LangChain (2026). smithtune v0.1.0 [software]. github.com/langchain-ai/smithtune. Accessed September 28, 2026.
  4. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O. and Narasimhan, K. (2024). SWE-bench: Can language models resolve real-world GitHub issues? ICLR 2024. arXiv:2310.06770. SWE-bench Verified subset: OpenAI (2024), Introducing SWE-bench Verified, openai.com/index/introducing-swe-bench-verified.
  5. Lambert, N. (2025). Reinforcement Learning from Human Feedback, chapter Rejection Sampling. arXiv:2504.12501. Online at rlhfbook.com.
  6. Nebius (2024). SWE-agent-trajectories and SWE-bench-extra [datasets]. huggingface.co/datasets/nebius/SWE-agent-trajectories.
  7. whilehq (2026). deepagents-review-four-arms recipe, whileai SDK [code and data]. github.com/whilehq/whileai-sdk/tree/main/recipes/community/deepagents-review-four-arms.

Run it

pip install whileai
git clone https://github.com/whilehq/whileai-sdk
cd whileai-sdk/recipes/community/deepagents-review-four-arms
python run.py --dry-run   # the comparison and the selection rule on fixture rows, no keys
python run.py traces      # stock agent over the training reviews, traces to LangSmith
python pick.py            # the lean traces, written back as LangSmith feedback
python make_results.py    # paired comparisons and the noise floor