Recipe / September 28, 2026
Benchmaxxing ParseBench
We built an open-weight document parsing agent for LlamaIndex's ParseBench. It scores 78.9, first among open-weight systems. The gain came from the harness around the model, not training. The recipe is public so you can climb from here.
The short version. ParseBench scores how well a parser turns real PDF pages into text an agent can use. We wrapped an open model that scores 70.8 alone in a harness that calls it several times per page. It now scores 78.9, first among open-weight systems. We tuned on one fifth of the documents and held out the rest. Training on charts did not help.
The benchmark rewards structure, not text
ParseBench [1] grades about 2,000 pages on tables, charts turned back into data, faithful text, formatting, and whether each element is boxed in the right place (grounding). A chart summarized in prose earns nothing. A markdown table earns nothing. So we read the scorer code first and shaped every output to what it checks.
The harness does the work
The model reads each page once to find its parts. Then specialists take over. Tables are re-read from a close crop. Charts are re-read from the whole page, so titles and legends stay in view. Text is re-read at full resolution to catch bold and superscripts. A layout model draws the boxes.
Each gain holds at 95% confidence. Voting on charts added nothing measurable.
| Item | Value |
|---|---|
| Table and chart specialists | 3.6 |
| Layout model | 2.8 |
| Styling from crops | 2.6 |
Our first version lost points on charts. Cropping cut off titles and legends, where the scored labels live.
Where it lands
First among open-weight systems. Within noise of Opus 5.5.
| Item | Value |
|---|---|
| LlamaParse Agentic | 87.0 |
| Opus 5.5 | 79.9 |
| while.ai | 78.9 |
| rakedoc-nano | 77.2 |
| Qwen3.8-27B alone | 70.8 |
Held out means held out
ParseBench has no training split, and we never trained on its pages. We split it by source report: one in five for tuning, the rest held out.
Some tuning did not carry over to unseen reports.
| Item | Value |
|---|---|
| Tuned | 80.0 |
| Held out | 77.1 |
Training the chart reader did not help yet
We trained the model on synthetic charts with reinforcement learning, rewarded by ParseBench's own chart check [3][4]. It learned those charts, not ParseBench's.
Training did not help.
| Item | Value |
|---|---|
| No training | 86.7 |
| Trained, single charts | 82.1 |
| Trained, chart pages | 85.9 |
Where to climb next
Tables are the cheapest points left.
| Item | Value |
|---|---|
| Charts | +9.9 |
| Grounding | +11.2 |
| Tables | -13.8 |
| Faithful text | -5.4 |
| Formatting | -6.7 |
For researchers
| What | Setting |
|---|---|
| Model | Qwen3.8-27B, vLLM with the qwen3 reasoning parser |
| Layout boxes | PP-DocLayoutV3 [2], score at least 0.2 |
| Chart pass | whole page at 2048 px, medium thinking, 3 samples, per-cell median |
| Table pass | crop, medium thinking, 3 samples, per-cell vote |
| Split | SHA-1 of source report name, 1 in 5 to tuning |
| Bands | percentile bootstrap over documents, 2,000 draws |
| Chart RL | CISPO [4], LoRA rank 32, group 16, ChartNet [3] minus overlap with any ParseBench chart |
| 95% bands | full 78.9 (77.8 to 79.9); held out 78.5 (77.3 to 79.6); step gains 3.6 (1.0 to 6.2), 2.8 (1.4 to 4.2), 2.6 (0.9 to 4.2); chart vote 0.2 (-1.0 to 1.4) |
Run it
The recipe runs everything on Modal with your own keys and reproduces every number here: github.com/whilehq/whileai-sdk/tree/main/recipes/04-train/parsebench.
References
- Zhang et al. ParseBench: A Document Parsing Benchmark for AI Agents. arXiv:2604.08538, 2026.
- PaddlePaddle. PP-DocLayoutV3. huggingface.co/PaddlePaddle/PP-DocLayoutV3_safetensors.
- IBM Granite. ChartNet, core_permissive subset. huggingface.co/datasets/ibm-granite/ChartNet.
- Khatri et al. The Art of Scaling Reinforcement Learning Compute for LLMs (ScaleRL, CISPO). arXiv:2510.13786, 2025.