Nutrient benchmark
Form Field VLM Leaderboard
End-to-end form-field extraction on 100 clean held-out pages with 701 annotated fields. Systems are ranked by Typed F1 at IoU 0.5: a field needs both a sufficiently tight box and the correct fine field type.
How to read it. Cells show IoU 0.5 / IoU 0.2. IoU 0.5 is the strict production-oriented threshold and determines ranking. IoU 0.2 diagnoses approximately correct but loose localization. Typed F1 additionally requires the correct fine field type. Box F1 evaluates localization while penalizing both missed and extra boxes. Box Recall measures field coverage without penalizing extra boxes. Cloud names identify the model generation evaluated in this benchmark; Details retain the exact submitted API model ID. Local seconds/page are comparable only within the recorded batch environment; provider batch APIs do not expose comparable request latency. Open any row for precision, recall, field-type, density, count, and grounding details.
| Rank | System | Family | Typed F1 .5 / .2 | Box F1 .5 / .2 | Box Recall .5 / .2 | sec/page |
|---|
open downloadable or openly documented weights · commercial Nutrient commercial weights · provider hosted API.
Submit or reproduce a result
Generate predictions for the frozen benchmark, score them with the reference scorer, and submit the resulting JSON for review.
python score.py --benchmark-repo nutrientdocs/form-field-vlm-benchmark \ --predictions predictions.json --out result.json
The benchmark and scoring contract are public; the fixed 100-page set will not change silently.
Models benchmarked: nutrientdocs/form-field-vlm · Qwen/Qwen3.5-4B · Qwen/Qwen3-VL-4B-Instruct · google/gemma-3-4b-it · google/gemma-4-E4B-it · mistral-community/pixtral-12b