Blog

Recipe / September 28, 2026

Benchmaxxing ParseBench

We built an open-weight document parsing agent for LlamaIndex's ParseBench. It scores 78.9, first among open-weight systems. The gain came from the harness around the model, not training. The recipe is public so you can climb from here.

The short version. ParseBench scores how well a parser turns real PDF pages into text an agent can use. We wrapped an open model that scores 70.8 alone in a harness that calls it several times per page. It now scores 78.9, first among open-weight systems. We tuned on one fifth of the documents and held out the rest. Training on charts did not help.

The benchmark rewards structure, not text

ParseBench [1] grades about 2,000 pages on tables, charts turned back into data, faithful text, formatting, and whether each element is boxed in the right place (grounding). A chart summarized in prose earns nothing. A markdown table earns nothing. So we read the scorer code first and shaped every output to what it checks.

The harness does the work

The model reads each page once to find its parts. Then specialists take over. Tables are re-read from a close crop. Charts are re-read from the whole page, so titles and legends stay in view. Text is re-read at full resolution to catch bold and superscripts. A layout model draws the boxes.

Points each step added

Each gain holds at 95% confidence. Voting on charts added nothing measurable.

Points each step added
ItemValue
Table and chart specialists3.6
Layout model2.8
Styling from crops2.6

Our first version lost points on charts. Cropping cut off titles and legends, where the scored labels live.

Where it lands

ParseBench overall, full set

First among open-weight systems. Within noise of Opus 5.5.

ParseBench overall, full set
ItemValue
LlamaParse Agentic87.0
Opus 5.579.9
while.ai78.9
rakedoc-nano77.2
Qwen3.8-27B alone70.8

Held out means held out

ParseBench has no training split, and we never trained on its pages. We split it by source report: one in five for tuning, the rest held out.

Tuned vs held-out documents

Some tuning did not carry over to unseen reports.

Tuned vs held-out documents
ItemValue
Tuned80.0
Held out77.1

Training the chart reader did not help yet

We trained the model on synthetic charts with reinforcement learning, rewarded by ParseBench's own chart check [3][4]. It learned those charts, not ParseBench's.

Chart score on the tuning documents

Training did not help.

Chart score on the tuning documents
ItemValue
No training86.7
Trained, single charts82.1
Trained, chart pages85.9

Where to climb next

while.ai vs Opus 5.5, by dimension

Tables are the cheapest points left.

while.ai vs Opus 5.5, by dimension
ItemValue
Charts+9.9
Grounding+11.2
Tables-13.8
Faithful text-5.4
Formatting-6.7

For researchers

WhatSetting
ModelQwen3.8-27B, vLLM with the qwen3 reasoning parser
Layout boxesPP-DocLayoutV3 [2], score at least 0.2
Chart passwhole page at 2048 px, medium thinking, 3 samples, per-cell median
Table passcrop, medium thinking, 3 samples, per-cell vote
SplitSHA-1 of source report name, 1 in 5 to tuning
Bandspercentile bootstrap over documents, 2,000 draws
Chart RLCISPO [4], LoRA rank 32, group 16, ChartNet [3] minus overlap with any ParseBench chart
95% bandsfull 78.9 (77.8 to 79.9); held out 78.5 (77.3 to 79.6); step gains 3.6 (1.0 to 6.2), 2.8 (1.4 to 4.2), 2.6 (0.9 to 4.2); chart vote 0.2 (-1.0 to 1.4)

Run it

The recipe runs everything on Modal with your own keys and reproduces every number here: github.com/whilehq/whileai-sdk/tree/main/recipes/04-train/parsebench.

References

  1. Zhang et al. ParseBench: A Document Parsing Benchmark for AI Agents. arXiv:2604.08538, 2026.
  2. PaddlePaddle. PP-DocLayoutV3. huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors.
  3. IBM Granite. ChartNet, core_permissive subset. huggingface.co/datasets/ibm-granite/ChartNet.
  4. Khatri et al. The Art of Scaling Reinforcement Learning Compute for LLMs (ScaleRL, CISPO). arXiv:2510.13786, 2025.