GDA: personal data out of your documents, on your own servers
Upload a contract, a decision or a case file. GDA finds every name, tax number, address and identifier, you review the findings in the browser, and it writes a copy in the same format with consistent pseudonyms. Nothing leaves your network.
The Contractor, Ioannis Markidis, tax no. 123456789, residing at 45 Knossou Ave., Heraklion, tel. 6912345678, undertakes to deliver the services described in Annex A within 30 days.
The Contractor, PER_1, tax no. AFM_1, residing at ADDRESS_1, tel. PHONE_1, undertakes to deliver the services described in Annex A within 30 days.
Illustrative example with fictitious data.
Extract
Text with positions: the PDF text layer, OCR for scanned pages, the XML of a Word file including headers, footers, footnotes and comments.
Detect
Validated patterns for ΑΦΜ, ΑΜΚΑ, IBAN, cards, phones and e-mails. A multilingual entity model plus Greek form heuristics for names, addresses and dates. Your own term lists.
Verify
A local language model labels every candidate: person, home address, birth date, organisation, place or nothing. Validated identifiers skip it.
Review
In the browser, on the rendered page: accept, reject, pick a pseudonym, add what was missed. One decision covers every occurrence of the same entity.
Apply
Positional edits in the original format. Scanned pages are re-rendered with the redactions burned in, never overlaid.
Built for Greek documents, not adapted to them
Consistent pseudonyms
The same person is PER_1 on page 1 and on page 48, the same tax number is AFM_1 throughout. The redacted copy stays readable and the relationships between the parties survive.
Greek identifiers, validated
ΑΦΜ, ΑΜΚΑ, IBAN and card numbers are checked against their checksums, Greek phone numbers are recognised in every spacing, identity cards and passports by their patterns. Company tax numbers are told apart from personal ones.
Names, addresses and birth dates
A multilingual entity model and Greek form heuristics find names in any grammatical case, patronymics, abbreviated initials and street addresses. The language model separates birth dates from contract dates and home addresses from company seats.
Scanned pages included
Pages without a text layer go through OCR, and the output page is replaced by an image with the redactions burned in. A digital signature stamp over a scan does not fool it.
Word documents, edited in place
Headers, footers, footnotes and comments are covered, tracked changes are accepted first, and the result is still an editable DOCX.
Review before anything leaves
Every finding is drawn on the rendered page. Accept, reject, retype or add missed items, with keyboard shortcuts. Findings left undecided are redacted, so a hurried reviewer cannot leak by omission.
Profiles and term lists
Switch entity types on or off, keep always-redact and never-redact lists (your own company's contact details, for instance) and save them as named profiles for different document families.
An API integrators can build on
Upload, poll, read and edit findings, apply, download: the same REST API the interface uses, with an API key, documented in OpenAPI and served offline.
Your documents never leave the building
GDA is delivered as a signed, offline installation bundle for your own server. Every component that touches a document, from text extraction to the language model, runs inside that installation.
Where documents must be shared, but people must not be
Universities and research institutes
Real administrative, legal or financial records used for research or teaching, anonymised once with a documented review and reusable by the whole team.
Courts, lawyers and notaries
Decisions, deeds and case files shared for publication, research or training, with parties and witnesses pseudonymised consistently.
HR, payroll and procurement files
Contracts, payslips and offers handed to auditors, consultants or bidders without exposing employees and representatives.
Documents for AI and analytics
Build training sets, retrieval corpora and evaluation sets from real documents, and feed GRAG or GStT output onward, without putting people into a model.
Access-to-documents requests
Release the document, not the third parties in it.
Banks, insurers and telecoms
Loan files, claims and customer correspondence passed between departments and vendors under data minimisation.
One license per installation, nothing per user or per document
GDA is licensed per installation for a fixed term, with unlimited users and documents. We install it on your server from the signed bundle, or your team does with the included installer and runbooks. Upgrades, backups and restores are scripted.
Answers before you ask
Does any document, or any part of one, leave our network?
No. Text extraction, OCR, the entity model and the language model run inside the installation on your server, on internal networks with no route out. There is no telemetry, no update check and no licence call home. The security pack explains how your own team can verify each of these.
Which formats does GDA handle?
Plain text, PDF with a text layer or scanned, and Word DOCX. The redacted copy comes back in the same format. Images and legacy .doc files are not accepted in this version.
Which personal data does it detect?
Names, home addresses, birth dates, ΑΦΜ, ΑΜΚΑ, IBAN, card numbers, Greek phone numbers, e-mail addresses, identity card and passport numbers, and optionally organisations, places and company tax numbers. Your own term lists add anything specific to you.
What hardware do we need?
One server of your own, physical or virtual, running Linux. Ordinary processors are enough; a graphics card makes the verification step faster. No cloud account and no internet connection are needed, and we size the machine with you during setup.
Can it run without the language model?
Yes. Verification can be switched off per profile. Detection then relies on patterns, the entity model and heuristics, runs faster, and errs towards redacting more. On the measured sets the language model's main effect is fewer false alarms, not more findings.
Who can use it, and what is logged?
Named users with two roles: administrators manage users, profiles and settings; reviewers upload, review, apply and download. Every action is in the audit trail with the acting user. Application logs never contain document text, personal data or client addresses.
Is pseudonymisation enough for the GDPR?
Pseudonymised data remains personal data under Article 4(5) of the GDPR. GDA is a measure of data protection by design and by default (Article 25) and of data minimisation, not an exit from the Regulation. Whether a given release is lawful remains your assessment; GDA gives you a consistent, reviewed and auditable way to make it.
See GDA on your own documents
Book a demo and we walk you through detection, review and output on sample documents. If it fits, we set up a 30-day evaluation on your own server.
Request a demo