Measured accuracy
How filterit measures itself
For: data protection officers, auditors, security reviewers and engineers evaluating the product. Scope: the method. The numbers it produces are on the accuracy page, each with the version of this method it was measured under.
A vendor's accuracy figure is worth exactly as much as the method behind it. A method you cannot inspect is a number you should not trust, including ours. This document is written so that a competent reviewer can find our thumb on the scale if it is there.
1. What the product is, legally
filterit applies pseudonymisation as a technical measure under Article 32(1)(a) GDPR and supports data protection by design under Article 25. It is not anonymisation: the mapping from each placeholder to its original value is kept, encrypted, because restoring the model's answer is the product.
Pseudonymised data remains personal data. Recital 26 and Article 4(5) GDPR are explicit: data that can be attributed to a person with additional information, and filterit deliberately keeps that information, stays within the scope of the Regulation.
Using filterit does not take a controller outside the GDPR, does not remove the need for a lawful basis, does not remove transparency obligations and does not by itself discharge a DPIA. It is a security and data minimisation control that reduces what a third-party model provider receives. That is a real and useful thing. It is not an exemption.
The measurements below quantify how well that control performs on our evaluation sets. They do not certify compliance, because no measurement can.
2. What we measure
Three programmes, reported separately and never averaged into one "accuracy" figure, because they fail in different ways and a customer needs to know which one is failing.
| Programme | Question it answers | Status |
|---|---|---|
| Detection quality | Of the personal data present, how much is found, and how much ordinary text is wrongly covered? | Measured, published per release on the accuracy page |
| Reversibility correctness | Is masking reversible without loss, and can a placeholder ever resolve to the wrong value? | Enforced as a build gate (section 6) |
| Answer fidelity | Does masking degrade the answer the model gives back? | Method drafted, not yet measured; nothing is claimed |
3. Detection quality
3.1 The unit of measurement
A span: a contiguous run of characters carrying one piece of personal data, with a type (NAME, AFM, IBAN and so on). Predicted spans are scored against spans annotated by a person.
All text is normalised to Unicode NFC before any offset is computed, on both sides. Greek is heavily accented, and the same name in decomposed form shifts every subsequent offset by one per accent; comparing offsets across normalisation forms produces silent, plausible-looking nonsense. Offsets are Unicode code point indices, half open.
3.2 Three ways to score one prediction
We publish three figures per set and per detection level, and we say which one the release gate reads.
| Figure | A prediction counts when it has | What it tells you |
|---|---|---|
| EXACT | the right type and identical boundaries | The quality number: the placeholder covers exactly the value, nothing more, nothing less |
| PARTIAL | the right type and any character overlap with the annotated span | The coverage number: the region was treated as that kind of data |
| Leak recall | any type, with the annotated span fully inside the prediction | The safety number: did anything get through uncovered? |
The release gate reads PARTIAL recall, and reads it per type: at the deep level, one annotated span of any type in either gold set that no prediction of that type touches fails the build. PARTIAL precision has a floor too (0.90 on the Greek set, 0.95 on the English set), and on the Greek set every type with annotated spans must hold EXACT F1 of at least 0.90.
Leak recall is reported next to them because the gap between leak recall and typed recall is entirely type confusion: data that was covered but labelled as something else. That matters for the audit trail and not at all for what the model provider receives, and collapsing the two into one figure would misrepresent whichever one the reader cares about.
Why the leak figure is strict about containment rather than overlap: filterit does not identify personal data, it replaces it. Every annotated character a prediction fails to cover is a character that stays in the text sent to the model provider. Under an 80 percent overlap rule, a nine digit Greek tax number with eight digits masked scores as a success. ΑΦΜ: [AFM_1]3 is not a success, and we do not want a metric that says it is.
Two consequences, both reported:
- Merged spans count in full. One prediction covering two adjacent names of the same type counts as two correct detections. The region is covered; nothing leaked. A stricter one to one matching rule would understate the product, and misstating in either direction is a defect in the metric.
- Boundary misses are separated from blind misses. A prediction with the right type that overlaps but does not fully contain the annotated span is counted as a miss and as a false positive, in its own category. It is a different failure from seeing nothing at all, and merging the two hides which is happening.
3.3 False positives, by kind
Not all over-masking costs the same, so it is not aggregated into one count.
| Kind | Definition | Cost to the customer |
|---|---|---|
| Type confusion | Right region, wrong type | Audit trail accuracy only. Nothing leaked, nothing lost |
| Boundary | Right type, incomplete coverage | Residual characters exposed and a slightly mangled prompt |
| Spurious | No personal data there at all | Pure utility loss: ordinary text replaced by a placeholder |
Alongside the raw count we compute a utility weighted false positive rate: each false positive is weighted by how many times the affected term occurs in the document (capped at ten) and by whether it falls inside the user's own instruction rather than the attached context. Covering a term that appears once in an annex costs almost nothing; covering one that appears twelve times rewrites the document's vocabulary.
This weighting is a hypothesis about harm, not a measurement of it. The factors are plausible and auditable, but nothing yet shows they predict real answer degradation. Testing that against measured answer fidelity is planned; if the correlation does not hold, the scheme gets replaced rather than defended.
3.4 Two detection levels, both published
The product runs two depths of the same pipeline, and a number is meaningless without saying which one produced it.
| Level | Where it runs | What it contains |
|---|---|---|
| Standard | The chat, where the answer must come back fast | The deterministic tier (format and check digit validators, identifier lists, name lists and surname morphology, context anchors, address grammar) and the statistical named entity recogniser |
| Deep | Documents, scanned files, visual redaction, the review grid, meetings | Everything in the standard level, plus the transformer models for names and for Greek entities |
The release gate reads the deep level. The standard level is published next to it, informational, so a chat user can see what the level they are using achieves rather than the best number the product can produce.
3.5 Operating point: precision is never quoted alone
Precision and recall trade against each other, so a precision figure without the recall it was measured at is not information. Every figure we publish carries its configuration: the set, its version, the detection level, the replay mode and the commit the numbers were produced from. The published configuration is the one customers run.
Part of the detection stack is deterministic: check digit and format rules with no confidence threshold. They cannot be tuned down, so they set a floor on recall and a ceiling on precision that no threshold sweep can cross. Where a configuration cannot reach a target, we report what it reaches. We do not interpolate to a rounder number.
3.6 Attribution across detection tiers
The stack has four tiers: deterministic rules, a statistical recogniser, transformer models (deep level only) and an optional verifier that ships switched off. Where several tiers detect the same span, it is attributed to the earliest tier that would have sufficed alone, because the question worth answering is "which tiers could be switched off without losing this?", and crediting the most sophisticated component for spans a rule already had would make every subsequent decision wrong.
For false positives the same ordering applies but attribution is non exclusive: every tier that contributed is counted, so per tier false positive counts sum to more than the total. That is deliberate and labelled wherever it appears.
4. The data
4.1 Two kinds of set, opposite purposes
Gold sets: annotated documents containing known personal data. They measure recall. The Greek set covers names across grammatical cases, gendered surnames, capitals, lower case, Latin transliterated Greek, greeklish, code switched lines and the structured Greek identifiers (ΑΦΜ, ΑΜΚΑ, identity card, IBAN, plates, phones). The English set covers names with honorifics, initials, particles and hyphens, addresses, and the identifiers of the English speaking countries. Both are written by hand, both are held out: no detector is tuned on them.
Adversarial negative set: documents that contain no personal data but are constructed to look as though they do. It measures precision. Six trap families: legal citations, organisation names that are also surnames, public classification codes, check digit look-alikes, greeklish variants and technical identifiers. Every detection in these documents is a false positive by construction, with one exception recorded per trap: a company named after a surname is correctly a company, and only the person label is the error that trap exists to catch.
The negative set matters more than its size suggests. Without it, precision would be measured only on whatever incidental false positives the gold sets happen to contain, which is a sample size nobody should build a claim on.
4.2 How the negative set is built, and why its numbers can be trusted
Fully synthetic, from a deterministic generator kept with the set. Re-running it reproduces the set byte for byte, and the continuous integration build does exactly that on every change, so a difference means the generator changed, not that we regenerated until the numbers improved.
Check digit traps are validated against the real algorithms at build time: a "negative" that turns out to be a genuinely valid identifier fails the build, because one mislabelled row poisons the whole precision figure. A deliberate subset does satisfy the check digit (protocol numbers that pass the ΑΦΜ check by construction), because those can only be rejected from context, and we want that measured rather than assumed.
No real document was copied. No real person, company, address, identifier or account number appears.
4.3 The rule that keeps the negative set meaningful
No detector is tuned against it. The moment a rule is adjusted to pass a specific trap, the set stops measuring precision and starts measuring memorisation.
Enforced procedurally: ten percent is held out and reported separately, aggregate results are read per category rather than per row during development, and if a trap must become a training signal it is copied into a test fixture and deleted from the set, recorded in the commit. The set shrinks visibly rather than degrading silently. The held out split is published next to the main split so the reader can see whether the rules carry to documents they were not measured on.
4.4 Cases we annotate arguably
Some cases are policy questions, not technical ones, and we keep them out of every metric rather than quietly resolving them in whichever direction flatters the result. Currently isolated: references to public figures in case law, office holders identified by office rather than by name, and academic citations carrying real surnames. These are documented and excluded, and will be included only once the position is decided and stated.
4.5 Real documents
We also run the deterministic validators over a small set of real Greek public documents (stamps, scanned forms, labels under the values). Each finding becomes a permanent rule and a permanent test. Those documents cannot be shared, so their numbers are not published: an unverifiable number is not evidence. A pilot on your own documents is the way to see what the product does on them.
5. Reproducibility
| Property | How |
|---|---|
| No network in the measurement path | The responses of the model services are recorded once and replayed from committed fixtures. A measurement cannot be influenced by an upstream service having a good day |
| Deterministic | Fixed seeds, no clock reads inside measured paths, stable sort orders throughout |
| Versioned sets | Each set carries an explicit version. A breaking change creates a new version; the old one stays on disk so historical numbers stay reproducible |
| Committed floors | The release gate's floors live in a checked in file. Lowering one requires a recorded rationale in the commit that does it |
| One scoring implementation | Greek and English are scored by the same code, so a cross language comparison is not comparing two scorers |
| The published table cannot drift | The page you read is generated from the measurement output, and the build fails when the published file differs from a fresh run |
6. Reversibility correctness
Reversibility is treated as a hard invariant with a build gate, not as a metric with a target. It is verified by property based testing with tens of thousands of generated inputs covering accents in both normalisation forms, final sigma, capitals, mixed Greek and Latin words, greeklish, inflected surnames, structured identifiers, concurrent access and streamed answers split at every character boundary.
One result belongs in a public method document because it is counter intuitive:
Restoration is not always byte identical, by design. In Greek, a restored personal name is re-inflected to agree with its context: a name masked in the nominative is restored in the genitive after a genitive article. Without this the output would be broken Greek.
So the guarantee is precise: the restored text carries the same information, and a placeholder never resolves to a different person's value. It is not "the output is identical to the input", and we would rather say so than let a reader assume a stronger claim.
A coverage limit, stated because it is material: these properties are proven against the in memory and on device vaults. The database backed vault of the hosted service is asserted by design and by unit tests, not yet by property testing.
7. Evidence generation
The product can produce a compliance evidence pack from its audit trail. Two limitations we state rather than wait to be asked:
- The audit chain proves integrity and ordering, not absolute time. A hash chain shows rows have not been edited or reordered relative to each other. It does not prove when they were written. There is currently no trusted third party timestamp on the chain head, so the pack's time claims rest on the operator's clock.
- Seal strength varies by deployment. The on device pack carries a public key signature verifiable offline. The hosted pack is sealed with a symmetric authentication code, which means it is verifiable by the key holder, the operator, and not independently by a third party. That is a weaker property than "signed" suggests to most readers, and we do not let the word do work it has not earned.
8. What these measurements do not tell you
The most load bearing section of this document.
| Not covered | Why it matters to you |
|---|---|
| Performance on your documents | The sets are curated and synthetic. Real documents are messier: scanning noise, unusual layouts, domain vocabulary. Ask for a pilot on your own data |
| A guarantee of recall | No detection system finds everything. A high score on a set is not a promise about the next document, which is why the product shows you what it found before anything leaves and asks you to check it |
| Anything upstream of detection | If text recognition fails to read a scanned name, detection never sees it. Text recognition quality is a separate failure with a separate fix |
| Answer quality | Whether masking degrades the model's answer is a separate programme (section 2) and is not yet measured. Any claim about it is currently unsupported, including by us |
| What the model provider does with masked data | Out of our control and out of our measurement. filterit reduces what they receive; it says nothing about their handling of it |
| Sample size | The gold sets are in the low hundreds of documents. That is enough to catch regressions and not enough for a confident population level claim. Sample sizes are printed next to every figure |
| Compliance | These are engineering measurements. They can support an Article 32 argument. They cannot make it for you, and no vendor can |
9. What we publish and what we do not
Published: this method; the results of every release with their set, set version, detection level, sample size, gate floors and commit; the false positive counts of the negative set per level, main and held out; the fact and direction of any regression.
Not published, and why:
- A single headline accuracy number. Detection quality, reversibility and answer fidelity fail differently. One number would hide whichever is worst.
- A recall figure on its own. A recall of 1.000 on a curated set of a few dozen documents is true and misleading. It is always printed with its set, its sample size and its precision.
- The sets themselves. The gold sets and the negative set, and the generator, stay in the repository. The method is published so that a reviewer can apply it to their own documents, which is the measurement that matters to them.
- Results on real customer documents. Unverifiable by the reader and hazardous to the customer.
- Answer fidelity figures before the judging method is calibrated against human agreement. An automated judge that has not been shown to agree with a person is not a measurement.
- Precision without its recall, and any figure whose confidence interval we would have to hide to make it look good. Where the interval is wide, the interval is published.
10. Change control
This method is versioned. A change that could move a number (the matching rule, the taxonomy, set composition, the levels, the gate floors) increments the version and is recorded with a rationale. Results always state which version produced them, so a figure can never be quietly improved by redefining the measurement underneath it.
| Version | Date | Change |
|---|---|---|
| 1.0 | 2026-10-07 | First published version. Three scoring modes with the gate on PARTIAL recall; two detection levels published; the negative set reported per level with its held out split; the sets are not published |
Reviewers who find an error, an unstated assumption or a place where the method flatters the product are invited to say so through the contact on the security page. That is what publishing the method is for.