Model evidence · open method

Multiple sources. Visible trade-offs.

One benchmark cannot describe every task. This comparison gives each source the same space and includes missing and weaker results.

Published model evidence · five equally presented sources · checked 30 September 2026
Model profileDecision Index0.2.1Chance-corrected index / 100Independent communityWorkflow Evalsdatasets 2026-09-28 · code 0ac3b8aModel-reference agreementTypeSafe · reference vendorJev Rerank Bench2026-09-25Dataset-macro nDCG@10 / 1Independent authorJevBench1.5.1Composite score / 100Same founder as System1ImageJevBench0.1.4Composite score / 100Same founder as System1
s1-fastPlumb-4BNot measured¹Not measured¹Not measured¹71.56Not measured¹
s1-proWinnow-12B Q850.02Not measured¹0.63473.2343.56
TypeSafe referenceJev 1.13.057.9167.80.67072.13Not measured¹

¹ Not measured means no comparable published measurement was verified for this model; it is not a zero. Results concern the tested model implementations, not the System1 service. ImageJevBench uses a separate image path; image input for customers is coming soon.

Conflict of interest: JevBench and ImageJevBench are run by the founder of System1 Models. Workflow Evals is published by TypeSafe AI, the vendor of the reference model. It measures model-reference agreement, not human-labelled accuracy. No overall score is published while comparable coverage is incomplete. Method and sources →

Equal weight, visible gaps

All five benchmarks have equal-width columns and the same treatment. Metrics retain their own units: an index, reference agreement, nDCG and a composite score measure different things. Submetrics and model cards that repeat the same results are not counted as extra votes.

There is no overall score because a common set of comparable measured results is missing. Any future overall must use the unweighted arithmetic mean over a declared common benchmark set: mean(100 × score / scale). Missing values are never silently omitted or replaced with zero. Chance correction already applied by a source is retained.

Where our profiles do worse

s1-fast (Plumb-4B) has one published comparable result: its JevBench composite is below the TypeSafe reference. It has no comparable result on the other four sources.

The Winnow-Q8 result is below the TypeSafe reference on Decision Index and on retrieval using the same yes/no-per-pair method. Its JevBench composite is slightly higher, with overlapping confidence intervals; that metric includes assumed speed adjustments and estimated hosting costs. s1-fast’s intelligence axis is well below the reference, and its composite benefits from estimated speed and cost. These differences stay visible in the matrix. The image result does not establish an available customer image service.

Sources and limitations

Decision Index · 0.2.1

38 index benchmarks across five areas. Winnow Q8_0 is identified, but the tested checkpoint revision is not published. Model implementation results, not a System1 service run. The board uses its own label-lift runtime configuration; identical System1 engine configuration is not established.

Primary source → · Original data → · Published: 2026-09-28. Checked: 2026-09-30.

Workflow Evals · datasets 2026-09-28 · code 0ac3b8a

Four public vendor workflow datasets, 705 cases; unweighted mean of the headline consensus metric reported in each manifest for Jev 1.13.0: invoice and customer service use exact action sets (accuracy), security incidents use exact action sets (agreement), and agent traces use primary-action disposition. Reference mode: all questions, reasoning off. Reference labels blend large-model answers, with some fallback contributors; agreement is not human-labelled accuracy. Dataset manifests publish no Plumb/Winnow runs. A small hosted-profile measurement is being prepared and would not replace this full published cohort.

Primary source → · WorkflowEvals code → · Public datasets → · Published: 2026-09-28. Checked: 2026-09-30.

Jev Rerank Bench · 2026-09-25

Eight English retrieval datasets; 1,617 scored questions. Same yes/no-per-pair method for both models. Winnow Q8 checkpoint b2b1421 is explicitly identified. Its A5000 run is not System1 service performance. Public dataset training overlap cannot be excluded. nDCG is ranking utility, not classification accuracy. The reference also has a stronger multi-level, batched configuration; Winnow was not tested with that method, so the like-for-like row is not the reference model’s ceiling.

Primary source → · Published: 2026-09-25. Checked: 2026-09-30.

JevBench · 1.5.1

Official released edition shown above. Its composite weights intelligence, calibration, speed and cost equally. This comparison uses only published aggregate results; no sealed item text was used to build it. The Plumb and Winnow rows do not record tested checkpoint revisions. Their local-GPU speed is adjusted with an assumed multiplier and gateway allowance; their hosted costs are estimates, while the TypeSafe reference uses observed hosted API latency and list price. These axes do not establish System1 service latency or cost. Plumb intelligence is well below the reference and its composite benefits from estimated speed and cost. Winnow has a slightly higher composite than the reference, with overlapping confidence intervals.

Primary source → · Original data → · Published: 2026-09-29. Checked: 2026-09-30.

ImageJevBench · 0.1.4

Public edition shown above. Winnow Q8_0 uses an F16 vision projector in a separate image implementation. The row does not record the tested checkpoint revision or projector hash. Image input for System1 customers is coming soon. TypeSafe Jev 1.13.0 and Plumb have no comparable image rows; other image-model results are not substituted. The held next release is not used.

Primary source → · Published: 2026-09-29. Checked: 2026-09-30.

Other coverage and model cards

We also checked primary sources from Artificial Analysis, Epoch AI, Vals and Hugging Face. No verified exact Plumb/Winnow result was found in the sources checked. Backbone scores are not transferred to a fine-tune or hosted service. Author-published measurements remain labelled and do not count again as independent benchmark results.

Plumb pinned model card → · Winnow pinned model card → · Artificial Analysis → · Epoch AI → · Vals →

Exact values, check date, sources, hashes and revisions: source data and exact values.