Skip to content
Offra

Blog · Method

Why we don't just hand the specification to ChatGPT

How Offra turns hundreds of pages of tender documents into a list of requirements, each tied to its sentence, and why one request to an AI can't.

By , founder ·

A Quebec public tender is not a document. It is a file: the specification (the devis), the price schedule, the drawings, the forms, and then the addenda that arrive while the tender is open. Across the notices published on SEAO, Quebec's public tendering system, since June 2024, half of these files run to 123 pages or more, and one in ten runs past 500. Somewhere in there are the few sentences that decide whether your bid can be accepted at all: a guarantee to attach, a licence to hold, a mandatory site visit, a form to sign.

The compliance matrix is the list of those sentences. Every requirement in the file on its own row, with the document and clause it comes from, and a link that opens the document at the quoted sentence. Offra builds it from the documents you upload.

The question we are asked most often is a fair one: why not just give the specification to ChatGPT, or Claude, or any other assistant, and ask it for the list? We use language models for this work ourselves. The problem is not the model. It is asking it for everything at once.

The short version

Imagine being handed a 300-page specification and asked, in a single read, for a list of everything it requires. You would end up writing a summary. A good summary, perhaps, but a summary.

Now imagine a team. The specification is cut into bundles of about 17 pages, each one spilling slightly into the next so that no sentence falls into the gap. Each person reads their bundle and, for every obligation they find, fills in the same form: what has to be done, whether it is mandatory, when, and the exact sentence from the specification, copied word for word. A coordinator gathers the forms and removes the duplicates. Finally, a proofreader takes every copied sentence and looks for it in the original document, to check it is really there.

That is what Offra does, with a language model in place of each reader, and ordinary code in place of the coordinator and the proofreader.

  1. 01

    Read

    Every page becomes text, and every line keeps its place on the page.

  2. 02

    Split

    Passages of about 17 pages, overlapping slightly.

  3. 03

    Extract

    Each passage read on its own, on the same form, the exact sentence copied.

  4. 04

    Merge

    Duplicates removed. Between two readings, the stricter one wins.

  5. 05

    Check

    Every quote found again in the document, or flagged as not found.

Result: the compliance matrix (sample)

  • Provide a 10% bid bondDevis · art. 7.2Exact match · p. 14
  • Attend the mandatory site visitAddenda 1 · art. 2Exact match · p. 1
  • Hold a valid RBQ licenceDevis · art. 3.1Approximate match · p. 6
The five steps, in the order they run. The result rows are a sample, not a real tender.

The difference between the two approaches is not a matter of intelligence. It is a matter of method: what makes the result reliable is the splitting, the same form for everyone, and the check at the end. None of those three exists when you ask a chatbot one question.

Why a single request is not enough

It summarizes instead of listing

Give a language model a long document and ask it for "all the requirements", and it will do what a person in a hurry would do. It will group them. We saw this on an 86-page specification read in one go: the answer came back as a single requirement, which said, in substance, to comply with the specification. True, and useless.

Illustration — invented requirements

The whole specification, in one request

  • Comply with all requirements of the specificationall of it

True, and useless: nothing to attach, sign or pay is named.

1 requirement

Passage by passage

  • Hold a valid RBQ licencepassage 1
  • Attend the mandatory site visitpassage 1
  • Attach the Revenu Québec certificatepassage 2
  • Provide a 10% bid bondpassage 3
  • Show proof of $2M liability insurancepassage 4
  • Submit a work schedulepassage 5
  • Complete and sign the price schedulepassage 6
  • Sign the bid formpassage 6

8 requirements

An illustration, with invented requirements. The case behind it is real: an 86-page specification, handed to the model in one go, came back with a single requirement.

It stitches and it rounds

Asked to quote its sources, a model tends to glue two passages together with an ellipsis, or to reword slightly. The sentence looks exact and exists nowhere in the document. As for the page number it gives, it cannot know it: what it read was text, not pages.

The same question does not always get the same answer

On the same passage of a specification, three identical requests returned 0, then 1, then 34 requirements. The model was the same; what changed was the server that answered. With nothing around it to notice, two of those three answers would simply have made a document disappear from your list.

Addenda move dates

An addendum can push the closing date back, but it can also bring it forward, or cancel it. So the most recent date written in the file is not necessarily the one in effect, and a model reading everything in one block has no reason to tell the two apart.

It does not tell you what it did not read

This is the most serious one. A chatbot always answers with confidence, whether it read all 300 pages or only the first 40, and nothing in its answer tells you what is missing. A useful compliance matrix has to be able to say: this document could not be read in full, and this quote, we could not find.


For technical readers

What follows describes the system as it runs today. The models change: we choose them by measuring them on this exact task, with the same instructions and the same form as production, and we switch when another one does better. The vendors that process your documents are named in our privacy policy.

Reading: text that keeps its place

Every PDF goes through optical character recognition (OCR), even when it already contains text. The reason is practical: OCR returns each block's position on the page, and that position is what later lets us box the exact sentence. Word documents and Excel workbooks are read directly.

Before any model sees anything, the text is cleaned. Running headers and footers repeated on at least half the pages are removed, table-of-contents lines ("4. Insurance .......... 12") are dropped, and the document's headings are put back as headings. A badly cleaned table of contents means a duplicate requirement, or worse, a requirement whose quote points at the table of contents rather than at the rule.

Splitting: 17 pages, with overlap

A document over 40,000 characters, about 17 pages, is cut into passages of that size which overlap by 3,000 characters. Cuts fall at the end of a paragraph or a line, and the overlap means a requirement straddling two passages is read whole at least once.

Each passage carries a header with the document's name and outline, so section references stay consistent from one passage to the next; the model is told never to extract anything from that context block. Workbooks follow a different rule. They are read in one piece as long as they stay a reasonable size, so that a price schedule remains one requirement rather than one per sheet, and beyond that they are split by row ranges, repeating the header row.

Extracting: a form, not a conversation

Each passage goes to a separate model call, up to eight at a time. The model does not answer in free text: it is required to call a tool, exactly once, and fill in a form whose fields are fixed in advance. For each requirement: a short statement; a category (technical, administrative, financial or legal); mandatory or not; when it applies (at bid, after award or during the contract); the deliverable to attach, if any; the action expected of the bidder; whether missing it is disqualifying; and a verbatim quote of at most 200 characters. The same form collects dates and legal clauses.

Three instructions weigh more than the others, and each answers a mistake we have seen:

  • At the level of the deliverable, never the field. A price schedule is one requirement, not one per line. A bid bond is one requirement, and its conditions (validity period, amount, signature) go inside it.
  • The quote is one continuous piece, copied character for character, with no "[...]". A patched-together quote exists nowhere and cannot be found.
  • Not everything is decisive. The model ranks each requirement by the attention it deserves, and keeps "decisive" for what a firm could genuinely fail to provide: a bond, a specific licence, an insurance threshold. Before this instruction, the bid bond came fourth, behind a cover letter.

The form is validated in our code, not by the model's vendor. That is not a detail: until early September 2026, one missing value in one requirement rejected the whole call, and an entire document vanished from the matrix. Today, a malformed requirement is dropped and the others are kept, and a truncated answer is salvaged up to its last complete item. When the model does not say whether a requirement is mandatory, it is treated as mandatory, because a mandatory requirement shown as optional can cost a contract. Each passage's result is saved as soon as it arrives, so an interrupted analysis resumes where it stopped.

When the model fails, and it does

A passage that returns nothing is not a passage that succeeded. This is the sneakiest failure: a turn that calls no tool ends cleanly, and the document counts as analysed without having contributed anything. For us, a passage has succeeded only once a row has been saved.

Otherwise a ladder of retries kicks in: the primary model, then a fallback model from another vendor, then, if the passage is too heavy, splitting it into two halves analysed separately. Every document ends with a status (read in full, read in part, or not read), and that status is shown to you.

Two causes of failure deserve naming, because they appear in no demo:

  • Reasoning eats the answer. Recent models "think" before answering, and that thinking counts against the same length limit as the answer. On one analysis instrumented end to end, 80% of the tokens generated were thinking, and two turns hit the limit without writing a single requirement.
  • The same name is not the same server. An open model is often offered by dozens of hosts, behind an intermediary that picks one per request. The one we were measuring had 27, and nine of them claimed to honour the obligation to call the tool while ignoring it. That is where the 0, 1 and 34 requirements above came from. The fix is simple and unglamorous: restrict the list to the hosts that passed the measurement.

Merging: the stricter reading wins

Passages overlap, and a specification often repeats the same obligation: in the instructions to bidders, in the administrative clauses, in the form. So there are duplicates, and they have to be removed without losing information.

The first pass is code, document by document. Two requirements are twins when their quotes land on the same spot on the page, when the text of one contains the other, or when they share almost all their words. When two twins disagree, the stricter one wins: if either says mandatory, the merged row is mandatory; if either says disqualifying, so is the merged row.

Four readings of the same specification

  • Provide a 10% bid bondpassage 2
    mandatorydisqualifyingat bid
  • Bid bond: 10% of the total pricepassage 3
    mandatorydisqualifyingat bid
  • Attach the bid bond (10%)overlap 2-3
    mandatorydisqualifyingat bid
  • Bid security of 10%passage 5
    mandatorydisqualifyingat bid

One row

Provide a 10% bid bond

mandatorydisqualifyingat bid

Every flag carried by at least one reading is kept.

Legal notes: two different figures, two rows

Bid guarantee: 10%Devis · art. 7.2
Bid guarantee: 5%Addenda 2 · art. 1

An addendum that changes the percentage leaves both in front of you. Neither is silently removed.

Duplicates are recognised first by their quote, when two passages overlap on the page, then by their text. A last pass groups the ones that say the same thing in other words, and it can only shorten the list.

Then comes a consolidation pass over the whole file, given to a model, to group what code cannot see: "Provide the insurance certificate, see Annex 3" and Annex 3 itself. This pass has one important property: it can only group or drop, never add, and its proposals are checked before they are applied. If it fails, nothing is merged, and you simply see a few more duplicates.

Legal notes follow a harder rule still. No similarity measure works for them: in our examples, two notes saying the same thing were less alike than two notes stating different rules. On a five-document demo file, extraction produced 52 legal notes, including the same bid guarantee four times, and the old merging method removed only three. So they merge only when their titles carry the same words, and never when those titles quote different figures. A 10% guarantee and a 5% guarantee stay two rows, because that is exactly what an addendum that changes the percentage produces.

Checking: every quote is looked up again

Any page number the model might offer is thrown away: it read only a passage of text, with no page markers in it. Instead, every quote is looked up again in the document, in steps, from the strictest to the loosest:

  1. the exact sentence, in an OCR block (which gives the box) or on a page;
  2. each piece of a patched-together quote, separately;
  3. the quote's distinctive words, when the wording has changed slightly;
  4. the section the model cited;
  5. otherwise, nothing.

The result is shown as it is, in the app's own words: "Exact match", "Approximate match", "Could not locate this passage in the document." A failure is never presented as a success.

What the model returned

« Le soumissionnaire doit joindre à sa soumission une garantie de soumission équivalant à 10 % du prix total soumis. »

Section cited: art. 7.2

No page number comes from the model: it only saw a passage of text.

Devis — page 14

Le soumissionnaire doit joindre à sa soumission une garantie de soumission équivalant à 10 % du prix total soumis.

Three possible outcomes

  • Exact match — page 14

    The sentence is on the page, word for word, and is boxed there.

  • Approximate match — page 14. The wording differs slightly from the document.

    The distinctive words are there; the exact sentence is not.

  • Could not locate this passage in the document.

    The requirement stays on the list, marked as such.

This check proves the quote exists in the document. It does not prove that the requirement drawn from it is right, which is why every row opens the page, so you can see for yourself.

It is worth being precise about what this check proves. It proves the quote exists in the document. It does not prove the requirement drawn from it was understood correctly, and a requirement whose quote cannot be found is not deleted: it stays on the list, marked as such. That is why every row of the matrix opens the document at the passage it found. The final check is yours, and we make sure it takes one click.

Addenda

Uploading an addendum re-runs the analysis of the whole file, not just the new document. The closing date in effect is decided explicitly, allowing for the fact that an addendum can bring it forward or cancel it, and the date it replaced stays visible, struck through, rather than being erased. The progress you had already recorded on each requirement carries over from one analysis to the next.

One limitation, stated plainly: when an addendum changes a requirement, the new version appears as its own row, with the addendum as its source. We do not decide for you which one replaces the other.

What the matrix is not

It is not a certification. It tells you what the file requires and where it says so. It does not promise that a bid ticking every row will be found compliant.

It publishes no rate. You will not read a percentage of requirements found here, because nobody, us included, has measured one on tenders set aside for the purpose. A vendor who gives you one should be able to tell you what it was measured on.

It does not touch the price. The documents you upload do not move the report's price range. They tell you what the file contains, not what the market will cost.

It starts with your documents. On SEAO, the specification is sold to the bidder; what Offra reads on the day a tender is published is the notice. The document analysis page explains the difference.

What the matrix does do, it does verifiably: every requirement carries its sentence, every sentence carries its page, and what could not be found is said.

What comes next

Offra scores every public tender currently open in Quebec against your company, and generates the full report on the ones that matter.