Skip to content
Offra

Measured accuracy

How often we are right, and what we are not right about.

Every measurement on this page is taken on 12,076 tenders the model had never seen: it was trained up to 2024-12-31, calibrated on the period after that, then measured on what came next. None of them is a target or an industry average — they are results on contracts already awarded.

An 80% range that held the price 95% of the time would be as wrong as one that held it 60% of the time: it would simply be too wide to use. What you should be looking for here is coverage near its target, and the width it took to get there.

01 — Does the range hold the price

Two ranges are published and each is calibrated separately against its own target. Neither is derived from the other.

80% range — real coverage
79.4%
Target 80%. Median width 4.2×.
50% range — real coverage
49.4%
Target 50%. Median width 1.9×.
Median gap between the centre and the real price
34.2%
Which is precisely why that centre is never shown as an answer.
Real prices below the range / above it
9.9% / 10.7%
Almost symmetric: the range does not lean one way.

What the calibration bought

The model's raw quantiles under-cover badly. A single shift applied everywhere corrects the average but widens the range by a third. Splitting by band corrects the average and makes the range narrower than it started — that is the one that is published.

Method80% coverageMedian width
Raw quantiles, no calibration71.1%4.3×
One shift, applied everywhere78.5%5.4×
Band by band — what is published79.4%4.2×

The tight range follows the same path: 41.9% raw against 49.4% published, for a target of 50%.

02 — Band by band

A correct average coverage can hide two halves wrong in opposite directions. So the calibration is done per predicted-price band, and each band is checked separately.

The target we hold ourselves to is coverage between 75% and 85% in every band. One band misses it on this measurement, and it is marked below. We publish it rather than smoothing it away: an accuracy page that shows only its successes measures nothing.

Predicted-price bandTenders80% coverage (target 75%85%)50% coverage (target 45%55%)
0 — the cheapest contracts68979.2%48.0%
160178.2%46.8%
267779.2%52.9%
367776.7%45.6%
469075.9%49.3%
569376.3%48.5%
668282.3%47.1%
760582.3%51.7%
854173.6% — outside the target45.8%
9 — the most expensive42979.7%47.1%

Measured on the tenders whose range rests on the model alone. Where the buyer publishes an estimate, it is built differently and measured separately — that is the next section.

03 — Why the bands are predicted-price bands

Because it is the only one you know at the moment of deciding — and because banding on the real price mechanically produces bad numbers, whatever the model.

Group the tenders by the amount finally awarded and coverage collapses at both ends. That is not a defect of the model: conditioning on the outcome selects, in the low band, precisely the contracts that went for less than expected. Any predictor, even a perfect one, would read the same. The table is here so that nobody has to discover it elsewhere.

GroupingLowMiddleHigh
By predicted price — what you know80.8%78.9%80.0%
By real price — known only afterwards57.9%88.6%65.0%

04 — On which tenders

The page average mixes two very different situations. Which one applies to you depends on what the buyer published with the notice — and the report tells you which is yours.

By what the buyer published

The largest gap on this page, by a distance. The buyer's estimate is the strongest lever we have.

SituationTenders80% coverage50% coverageMedian gapWidth
The buyer published an estimate5,79280.6%50.6%20.2%2.4×
No estimate published6,28478.3%48.4%54.3%7.4×
Market description recovered11,86879.4%49.5%33.9%4.2×
No description20881.3%45.7%74.0%20.5×
Tender documents published11,53479.4%49.5%33.0%4.1×
No documents54279.2%47.4%72.5%23.5×

By the nature of the contract

Far steadier. The category moves coverage very little.

SituationTenders80% coverage50% coverageMedian gapWidth
Goods2,85580.8%50.3%30.1%3.7×
Services4,32178.4%49.0%36.3%4.7×
Works4,74079.5%49.1%34.6%4.2×

The “width” column is the quantiles before calibration: the band-calibrated width is not measured separately on each situation, and giving one for the other would be a correct number under a wrong label. Across the whole set, calibration tightens the median width from 4.3× to 4.2×.

A notice with no published estimate is therefore measured at around 54.3% median gap against 20.2% when there is one. The report shows which of those two situations is its own — the reliability indicator, beside the range, and not in a footnote.

05 — And if the range is too wide

A range that is correct but very wide is no use to you. It is the most well-founded objection to this method, so here is the full distribution rather than its median.

80% ranges whose top is under twice the bottom
0%
Zero, and that is not a rounding: at that level of certainty, a range that narrow does not exist in our data.
… under three times the bottom
42%
… under ten times the bottom
73%
So 27% of 80% ranges are wider than that.
50% ranges under twice
52%
The tight range is the one you decide on.
Range widthNarrowest 10%Lower quarterMedianUpper quarterWidest 10%
At 80%2.4×2.6×4.2×10.4×19.4×
At 50%1.5×1.5×1.9×3.4×4.9×

The report places your range inside this distribution rather than leaving you to guess, and it refuses to draw one when the notice does not carry enough information to do better than an order of magnitude. A wide range is an admission of uncertainty; a very wide one is no longer a measurement at all, and it is better to say so than to draw it.

What the report does with that width

On a notice picked at randomShareWhat you see
The range is published73.8%Both intervals, the buyer's estimate if there is one, and the price steps
Published, but flagged as wide24.5%The same, with the warning placed beside the chart rather than in a note
No range at all1.7%The comparable contracts and the reason, instead. Neither a range nor a price to win

That last row is the notices with no buyer estimate and no usable specifications. For those we have no measured accuracy regime at all, and it is that fact — rather than the width of the interval — that decides. Those notices do not necessarily have the widest range; some are narrower than the quarter flagged above. An interval that looks confident on a contract we do not know how to value is the dangerous case, not the one that looks wide.

06 — What it compares against

Two simple methods act as a permanent guard rail. The first is the one anybody can apply; the second beats us on its own ground, and that is why we use it.

MethodTenders80% coverageMedian gapMedian width
The contract family's quantiles — the obvious method12,07671.6%79.0%30.3×
Our range, on the same contracts12,07679.4%34.2%4.2×
The buyer's estimate, recalibrated — where there is one5,79280.8%19.9%2.7×

Our range is 7.2× narrower than the obvious method, at comparable coverage, and is wrong half as often.

The third row beats us, and we publish it for that reason. On the 48% of notices where the buyer published an estimate, a simple correction applied to that figure does better than our 285 variables. We do not ignore it: the range published on those notices is a blend of the two, weighted on held-out data. A model that refused to use better information because it was not its own would be a bad product.

07 — The report's other forecasts

Price is not the only thing the report forecasts. The others are measured the same way, on the same split.

Mean error on the number of bidders
1.9
Measured on 8,540 tenders.
Real bidders appearing in the ten names proposed
43%
44% excluding companies with no SEAO history, which nothing can predict.
Of ten names proposed, those who really file
14%
One or two in ten. The list is for recognising who circles this buyer, not for guessing the opening of the envelopes.
Gap between stated probabilities and real frequencies
0.6%
When the report says 30%, the observed frequency is close to 30%.

08 — Where these numbers come from

Every measurement comes from a single results file, produced by the same evaluation cited in each report sold. No figure on this page is entered by hand: they are generated from that file, and a test fails if one of them is hard-coded.

That is the only thing that makes these claims checkable. An accuracy page copied by hand outlives the measurement it describes, and that is exactly what happened here before this mechanism existed.

Split
trained up to 2024-12-31
Calibration
up to 2025-06-30
Measurement
after that date, 12,076 tenders
Population
open public tenders, or 21% of the corpus
Estimate published by the buyer
48% of notices
Market description recovered
98% of notices
Tender documents
96% of notices

What these measurements do not say: what the contract costs you, what your order book is worth, or what gets decided outside the public process. The limits are written out in full in the terms of use, and the method in detail is in the report's description.

Get started

The report carries these same measurements beside every figure, on your tender rather than on the corpus.