Nutrient benchmark

Form Field VLM Leaderboard

End-to-end form-field extraction on 100 clean held-out pages with 701 annotated fields. Systems are ranked by Typed F1 at IoU 0.5: a field needs both a sufficiently tight box and the correct fine field type.

How to read it. Cells show IoU 0.5 / IoU 0.2. IoU 0.5 is the strict production-oriented threshold and determines ranking. IoU 0.2 diagnoses approximately correct but loose localization. Typed F1 additionally requires the correct fine field type. Box F1 evaluates localization while penalizing both missed and extra boxes. Box Recall measures field coverage without penalizing extra boxes. Cloud names identify the model generation evaluated in this benchmark; Details retain the exact submitted API model ID. Local seconds/page are comparable only within the recorded batch environment; provider batch APIs do not expose comparable request latency. Open any row for precision, recall, field-type, density, count, and grounding details.
RankSystemFamilyTyped F1 .5 / .2Box F1 .5 / .2Box Recall .5 / .2sec/page

open downloadable or openly documented weights · commercial Nutrient commercial weights · provider hosted API.

Submit or reproduce a result

Generate predictions for the frozen benchmark, score them with the reference scorer, and submit the resulting JSON for review.

python score.py --benchmark-repo nutrientdocs/form-field-vlm-benchmark \
  --predictions predictions.json --out result.json

The benchmark and scoring contract are public; the fixed 100-page set will not change silently.

Models benchmarked: nutrientdocs/form-field-vlm · Qwen/Qwen3.5-4B · Qwen/Qwen3-VL-4B-Instruct · google/gemma-3-4b-it · google/gemma-4-E4B-it · mistral-community/pixtral-12b

Nutrient

About the author
This project is maintained and funded by Nutrient - The deterministic document infrastructure enterprises run their highest-stakes workflows on: replayable output, clear exceptions, and full audit trails on the messy, regulated documents where AI alone breaks.