Measured accuracy
We do not ask you to trust us. We show you the measurements.
Every detection change passes automated quality gates. If even one personal-data item in our evaluation sets escapes unmasked, that release does not ship.
- 34 — personal data categories
- 10 — languages with measured coverage
- 100% — recall on our own evaluation sets, not a guarantee for every document
- 1,300+ — automated checks on every release
The scorecard, as CI produces it
The numbers per entity
From our gold sets, in hermetic replay. Everything here is written by the quality gate itself, not by hand.
What the numbers mean, in plain words
- Recall: Of 100 personal data items actually present in the text, how many we detected. 1.000 means none slipped past us in the set.
- Precision: Of 100 items we masked, how many really were personal data rather than a false alarm. 1.000 means no needless masking.
- F1: The two combined into one number, so that a system that simply hides everything does not look good. 1.000 is perfect.
- Leak recall: How many of the items that had to be hidden were covered at least in part. It is the number we care about most: one uncovered item is a leak.
Every figure refers to our own evaluation sets, not to every possible document. Always review the findings before you send a text.
The other side of the same problem
Not masking what does not need masking
We would rather mask one item too many than miss one (the “recall-first” principle). That carries an obvious cost: a system that masks everything never misses anything, and is useless. In September 2026 we measured and reduced false findings, under the explicit condition that recall was not allowed to drop by a single value.
- False findings on an adversarial set: 241 → 92 (62% reduction)
- Documents with no false finding at all: 64.1% → 86.9% (+22.8 points)
- Recall on the Greek reference set: 1.000, unchanged at every step
- Overall F1 (micro): 0.994
- The byte-exact reference snapshot stayed identical throughout
- Two candidate rules were rejected because the quality gate priced them
Language coverage
Three tiers. For each one we say exactly what holds.
Production depth
Greek · English
Full detection with Greek-specialised models and deterministic validators: AFM checksum, AMKA Luhn, IBAN, land-registry codes, IDs, plates. Name inflection handling and identity linking across documents. Validated on real public documents, not only curated text.
Measured coverage of names and structured data
German · French · Spanish · Italian · Portuguese · Dutch · Polish · Romanian
Names via a multilingual model at 100% recall on our evaluation sets, including inflected Polish forms. Name lists and surname morphology, plus deterministic dates, postal codes, phone numbers, VAT and IBAN.
Other scripts: flagged before sending
Cyrillic · Arabic · Chinese and other scripts
Names written in Cyrillic, Arabic or Chinese script are not detected yet. When a text contains them, filterit flags it before anything leaves, so the decision stays with you. Script-independent items such as IBANs, emails, phone numbers and card numbers are detected as usual.
How we measure
The methodology is the product. These are its four instruments.
01 Quality gates that only tighten
Labelled evaluation sets per language and data type. The recall floor is 100%: a single leak in any set automatically blocks the release. Thresholds ratchet upward and are never silently relaxed.
02 Real-document validation
Curated sets are not enough. Reality has stamps, scanned forms and captions printed under the values. We test on 22 real Greek public documents with labelled data, and every finding becomes a permanent rule with a permanent test.
03 Deterministic first
Anything with structure is verified by proof, not statistics: the Greek tax number by its checksum, AMKA by Luhn, IBAN by mod-97. Neural models handle only what cannot be proven, such as names and addresses, and are wrapped in confirmation layers.
04 When we cannot, we say so
Unreadable handwriting? An unsupported script? The product does not pretend. It shows an explicit warning before anything leaves. Silent failure is the only truly dangerous failure mode in a privacy tool.
What gets detected
34 categories, 11 of them specific to Greek documents
Greek identifiers are validated (check digit, format, context) rather than matched as digit runs, and names are followed across grammatical cases. Tap a category for details.
| Tax ID (ΑΦΜ) | Greek tax identification number with modulo-11 check digit validation, not a bare 9-digit match. | check: modulo 11 |
|---|---|---|
| Social security (ΑΜΚΑ) | Greek social-security number with Luhn and birth-date validation. | check: Luhn |
| ID card number | Greek police ID card, old and new formats (letters and digits). | check: μορφή |
| Vehicle plates | Greek registration plates, restricted to the letters Greek plates actually use. | check: μορφή |
| Court case numbers (ΓΑΚ, ΕΑΚ, ΑΒΜ) | Greek court case and docket numbers, context-gated with NLP. | check: συμφραζόμενα |
| Business registry (ΓΕΜΗ) | General Commercial Registry number, 12 digits, context-gated to avoid clashing with other numbers. | check: συμφραζόμενα |
| Cadastre code (ΚΑΕΚ) | National Cadastre property code, 12-digit and full 18-digit forms, with prefecture-code validation. | check: κωδικός νομού |
| myDATA and EFKA registries | AADE myDATA MARK and invoice UID, EFKA employer and insured numbers (ΑΜΕ, ΑΜΟΕ). | check: συμφραζόμενα |
| Property references (ΑΤΑΚ, ΠΕΑ) | E9 property identifier and energy-performance certificate security numbers. | check: συμφραζόμενα |
| Tiresias, auctions, BIC | Greek credit-bureau references, e-auction codes and BIC/SWIFT codes. | check: συμφραζόμενα |
| Greek places | Greek gazetteer of cities, municipalities and regions, in every grammatical case. | gazetteer |
| IBAN | Bank accounts of every EU/EEA country with MOD-97 validation. | check: MOD-97 |
|---|---|---|
| EU and UK VAT numbers | VAT numbers of all 27 EU countries with their national checks, plus UK VAT. | check: VIES |
| National IDs (EU, US, UK) | DNI, PESEL, Codice Fiscale, NIR, SSN/ITIN/EIN, National Insurance and more. | check: εθνικοί έλεγχοι |
| Passports | Passport numbers, context-gated in 24 EU languages. | check: συμφραζόμενα |
| Driving licences | Driving-licence numbers, context-gated. | check: συμφραζόμενα |
| Phone numbers | Landlines and mobiles for the EU, UK, US/Canada and Switzerland, validated per country. | check: libphonenumber |
| Payment cards | Card numbers with Luhn, together with CVV and expiry. | check: Luhn |
| Chassis number (VIN) | 17-character vehicle identification number per ISO 3779, context-gated. | check: ISO 3779 |
| Digital signatures | E-signature identifiers: DocuSign envelope, Adobe Sign agreement, SHA-1/SHA-256 certificate thumbprints. | check: μορφή |
| Postal codes | Greek postcodes, US ZIP codes and UK postcodes. | check: συμφραζόμενα |
| Names across grammatical cases | Greek NER that follows inflection, greeklish spellings and links the forms of the same person. | NER |
|---|---|---|
| Companies and organisations | Company and institution names, Greek and international. | NER |
| Email addresses | Email addresses. | check: μορφή |
| Street addresses | Street, number and area, in Greek and Latin-script text. | NER |
| Dates | Birth dates and other identifying dates. | check: μορφή |
| Trademarks | Trademark names from a private list of your own clients. | gazetteer |
| Medical data | Conditions, medication, blood type, record, insurance and prescription numbers, ICD-10 codes. | NLP |
|---|---|---|
| Special categories | Religion, nationality, political and trade-union references, marital status, profession. | NLP |
| API keys and secrets | API keys, tokens and JWTs, private keys, connection strings, passwords, typed in the placeholder. | check: μορφή |
|---|---|---|
| Reference codes | Prefixed identifiers such as EMP-, INV-, CUST-, ORD-. | check: μορφή |
| MAC addresses | Network device identifiers. | check: μορφή |
| IP addresses | IPv4 and IPv6. | check: μορφή |
| Technical identifiers | File paths carrying a username and internal hostnames. | check: μορφή |
In scanned PDFs and images, signatures and stamps are covered by visual redaction, with a preview before it is applied.
How to read the measurements
The figures on this page come from our evaluation sets and from the real documents we test on, and describe how the system behaves on those. On a new document performance can differ. That is why the privacy panel shows what was masked in every conversation, visual redaction goes through a preview before it is applied, and the system tells you wherever it cannot read the text. The final decision is always yours.
Measurements last updated: September 2026
See it on your own document
Upload an invoice or a letter and see what gets masked, before anything leaves.