How it compares
How filterit differs from the alternatives you have.
An office that wants to use AI on client documents has three routes today: paste into the AI chat directly, run a detection library on its own, or put a pseudonymisation tool next to the AI. This table compares approaches, not companies. Everything it says about filterit links to a page where you can check it.
| filterit | AI chat directly (ChatGPT, Claude and others) | Open-source detection library, your own deployment | Pseudonymisation tool without AI | |
|---|---|---|---|---|
| Detected personal data does not reach the AI provider | Yes — Placeholders in place of the detected names, tax ids, ΑΜΚΑ. The answer comes back with the values on your side | No — What you paste leaves as it is | Partly, conditionally or product-dependent — Yes, if you wire it before every call and maintain it | Yes — That is its job |
| Greek identifiers with validation (ΑΦΜ modulo 11, ΑΜΚΑ Luhn, ΚΑΕΚ, ΓΕΜΗ, court case numbers, myDATA) | Yes — 11 of the 34 categories exist only for Greek documents, with check digits or context | No — Detects nothing, reads whatever you give it | Partly, conditionally or product-dependent — Whatever you write and maintain | Partly, conditionally or product-dependent — Depends on the product. Ask about check digits and Greek inflection |
| Names across Greek grammatical cases and in greeklish | Yes — Greek NER that follows inflection, a gazetteer, and linking of the forms of the same person | No — Not applicable | Partly, conditionally or product-dependent — General NER models. Greek needs your own work | Partly, conditionally or product-dependent — Depends on the product |
| Published accuracy per category, on every release | Yes — Accuracy page with precision, recall and F1 from held-out gold sets, gated in CI | Not applicable — Not applicable | Partly, conditionally or product-dependent — Whatever you measure on your own data | Partly, conditionally or product-dependent — Depends. Ask for the methodology and the gold sets |
| AI inside the same product, with the values restored in the answer | Yes — Chat, files and email in one subscription. Placeholders become names again in your account | Yes — That is the product, without the masking step | No — A library, not a product | No — You bring your own AI and restore by hand |
| Reversible placeholders and pseudonyms with an encrypted mapping | Yes — [NAME_1] stable within a conversation, pseudonyms with correct inflection, AES-256-GCM vault | No — Not applicable | Partly, conditionally or product-dependent — Operators exist. You keep the mapping yourself | Partly, conditionally or product-dependent — Usually yes. Depends on the product |
| Visual redaction of PDFs and scans, with a preview | Yes — Black boxes over values, signatures and stamps, three-tier OCR | No — Not applicable | Partly, conditionally or product-dependent — Separate module, your own integration | Partly, conditionally or product-dependent — Depends on the product |
| Organisation policy, roles and proof for the DPO | Yes — Mandatory types per organisation, roles, evidence pack, egress ledger, PII-free reports | Partly, conditionally or product-dependent — The provider’s admin controls, no personal-data detection | No — Your own code | Partly, conditionally or product-dependent — Depends on the product |
| Where detection runs | Partly, conditionally or product-dependent — Web app: on our servers in the EU (Hetzner, Germany). Desktop agent and IDE plugins: locally, on your device | No — At the provider, together with the full text | Yes — Wherever you deploy it | Partly, conditionally or product-dependent — Usually on the product’s server |
| AI model provider | Partly, conditionally or product-dependent — Default today: Anthropic (Claude API), receiving masked text only. Residency policy per organisation, an EU-based provider as an option, a local model in development | Partly, conditionally or product-dependent — The provider you chose, with the full text | Not applicable — Calls no model | Not applicable — Calls no model |
| Integrations | Partly, conditionally or product-dependent — Web, Masking API, Python and Node SDKs, MCP server, VS Code, JetBrains, browser extension, Word add-in, desktop agent for Windows | Yes — Browser, Office, mobile, from the provider | Partly, conditionally or product-dependent — Whatever you build | Partly, conditionally or product-dependent — Depends on the product |
| Cost to start | Yes — Free with a character allowance, Pro with the AI included | Partly, conditionally or product-dependent — Per-user subscription at the provider | Partly, conditionally or product-dependent — Free code, your own running and maintenance cost | Partly, conditionally or product-dependent — Usually a subscription, without the AI |
Legend
- Yes
- Partly, conditionally or product-dependent
- No
- Not applicable
What you can check yourself
What gets detected
34 categories, 11 of them specific to Greek documents
Greek identifiers are validated (check digit, format, context) rather than matched as digit runs, and names are followed across grammatical cases. Tap a category for details.
Greek-specific identifiers
- Tax ID (ΑΦΜ) — Greek tax identification number with modulo-11 check digit validation, not a bare 9-digit match. (check: modulo 11)
- Social security (ΑΜΚΑ) — Greek social-security number with Luhn and birth-date validation. (check: Luhn)
- ID card number — Greek police ID card, old and new formats (letters and digits). (check: μορφή)
- Vehicle plates — Greek registration plates, restricted to the letters Greek plates actually use. (check: μορφή)
- Court case numbers (ΓΑΚ, ΕΑΚ, ΑΒΜ) — Greek court case and docket numbers, context-gated with NLP. (check: συμφραζόμενα)
- Business registry (ΓΕΜΗ) — General Commercial Registry number, 12 digits, context-gated to avoid clashing with other numbers. (check: συμφραζόμενα)
- Cadastre code (ΚΑΕΚ) — National Cadastre property code, 12-digit and full 18-digit forms, with prefecture-code validation. (check: κωδικός νομού)
- myDATA and EFKA registries — AADE myDATA MARK and invoice UID, EFKA employer and insured numbers (ΑΜΕ, ΑΜΟΕ). (check: συμφραζόμενα)
- Property references (ΑΤΑΚ, ΠΕΑ) — E9 property identifier and energy-performance certificate security numbers. (check: συμφραζόμενα)
- Tiresias, auctions, BIC — Greek credit-bureau references, e-auction codes and BIC/SWIFT codes. (check: συμφραζόμενα)
- Greek places — Greek gazetteer of cities, municipalities and regions, in every grammatical case.
European and international
- IBAN — Bank accounts of every EU/EEA country with MOD-97 validation. (check: MOD-97)
- EU and UK VAT numbers — VAT numbers of all 27 EU countries with their national checks, plus UK VAT. (check: VIES)
- National IDs (EU, US, UK) — DNI, PESEL, Codice Fiscale, NIR, SSN/ITIN/EIN, National Insurance and more. (check: εθνικοί έλεγχοι)
- Passports — Passport numbers, context-gated in 24 EU languages. (check: συμφραζόμενα)
- Driving licences — Driving-licence numbers, context-gated. (check: συμφραζόμενα)
- Phone numbers — Landlines and mobiles for the EU, UK, US/Canada and Switzerland, validated per country. (check: libphonenumber)
- Payment cards — Card numbers with Luhn, together with CVV and expiry. (check: Luhn)
- Chassis number (VIN) — 17-character vehicle identification number per ISO 3779, context-gated. (check: ISO 3779)
- Digital signatures — E-signature identifiers: DocuSign envelope, Adobe Sign agreement, SHA-1/SHA-256 certificate thumbprints. (check: μορφή)
- Postal codes — Greek postcodes, US ZIP codes and UK postcodes. (check: συμφραζόμενα)
People and contact details
- Names across grammatical cases — Greek NER that follows inflection, greeklish spellings and links the forms of the same person.
- Companies and organisations — Company and institution names, Greek and international.
- Email addresses — Email addresses. (check: μορφή)
- Street addresses — Street, number and area, in Greek and Latin-script text.
- Dates — Birth dates and other identifying dates. (check: μορφή)
- Trademarks — Trademark names from a private list of your own clients.
Health and special categories
- Medical data — Conditions, medication, blood type, record, insurance and prescription numbers, ICD-10 codes.
- Special categories — Religion, nationality, political and trade-union references, marital status, profession.
Code, secrets and systems
- API keys and secrets — API keys, tokens and JWTs, private keys, connection strings, passwords, typed in the placeholder. (check: μορφή)
- Reference codes — Prefixed identifiers such as EMP-, INV-, CUST-, ORD-. (check: μορφή)
- MAC addresses — Network device identifiers. (check: μορφή)
- IP addresses — IPv4 and IPv6. (check: μορφή)
- Technical identifiers — File paths carrying a username and internal hostnames. (check: μορφή)
In scanned PDFs and images, signatures and stamps are covered by visual redaction, with a preview before it is applied.
What we do not claim
We do not compare against named vendors: their features change and we do not measure them. There are tools with more languages and more integrations than ours. Our focus is Greek documents, Greek identifiers with validation, measured accuracy and AI inside the same product. No detection is infallible: that is why we publish accuracy per category and show the findings before anything leaves, so you can check and add what is missing. The masking is pseudonymisation, not anonymisation: the mapping exists, encrypted, so the values can return in the answer. The default model is Anthropic’s and receives the text with the detected details masked. Local-model and dedicated-server editions are in development.
The columns describe categories of approaches, not products. If you represent a product and believe a description applies to you and is inaccurate, write to contact@filterit.app and we will correct it.
See the difference on your own document
Paste an email or upload a contract and see what we find, before anything leaves.