GOVADI DATA ANONYMIZATION

GDA: personal data out of your documents, on your own servers

Upload a contract, a decision or a case file. GDA finds every name, tax number, address and identifier, you review the findings in the browser, and it writes a copy in the same format with consistent pseudonyms. Nothing leaves your network.

On-premises · Greek and English · PDF, Word and text · No internet connection required
REVIEW · SYMVASI_2026_0412.PDF · 6 PAGES
ON-PREM

The Contractor, Ioannis Markidis, tax no. 123456789, residing at 45 Knossou Ave., Heraklion, tel. 6912345678, undertakes to deliver the services described in Annex A within 30 days.

FINDINGS · 4 OF 23
Ioannis MarkidisPERPER_1
123456789AFMAFM_1
45 Knossou Ave., HeraklionADDRESSADDRESS_1
6912345678PHONEPHONE_1
REDACTED COPY

The Contractor, PER_1, tax no. AFM_1, residing at ADDRESS_1, tel. PHONE_1, undertakes to deliver the services described in Annex A within 30 days.

Illustrative example with fictitious data.

3
Formats in, the same 3 out: PDF, Word, text
14
Types of personal data detected
0
Bytes leave your network
AES-256
Encryption of every stored file
1

Extract

Text with positions: the PDF text layer, OCR for scanned pages, the XML of a Word file including headers, footers, footnotes and comments.

2

Detect

Validated patterns for ΑΦΜ, ΑΜΚΑ, IBAN, cards, phones and e-mails. A multilingual entity model plus Greek form heuristics for names, addresses and dates. Your own term lists.

3

Verify

A local language model labels every candidate: person, home address, birth date, organisation, place or nothing. Validated identifiers skip it.

4

Review

In the browser, on the rendered page: accept, reject, pick a pseudonym, add what was missed. One decision covers every occurrence of the same entity.

5

Apply

Positional edits in the original format. Scanned pages are re-rendered with the redactions burned in, never overlaid.

CAPABILITIES

Built for Greek documents, not adapted to them

01

Consistent pseudonyms

The same person is PER_1 on page 1 and on page 48, the same tax number is AFM_1 throughout. The redacted copy stays readable and the relationships between the parties survive.

02

Greek identifiers, validated

ΑΦΜ, ΑΜΚΑ, IBAN and card numbers are checked against their checksums, Greek phone numbers are recognised in every spacing, identity cards and passports by their patterns. Company tax numbers are told apart from personal ones.

03

Names, addresses and birth dates

A multilingual entity model and Greek form heuristics find names in any grammatical case, patronymics, abbreviated initials and street addresses. The language model separates birth dates from contract dates and home addresses from company seats.

04

Scanned pages included

Pages without a text layer go through OCR, and the output page is replaced by an image with the redactions burned in. A digital signature stamp over a scan does not fool it.

05

Word documents, edited in place

Headers, footers, footnotes and comments are covered, tracked changes are accepted first, and the result is still an editable DOCX.

06

Review before anything leaves

Every finding is drawn on the rendered page. Accept, reject, retype or add missed items, with keyboard shortcuts. Findings left undecided are redacted, so a hurried reviewer cannot leak by omission.

07

Profiles and term lists

Switch entity types on or off, keep always-redact and never-redact lists (your own company's contact details, for instance) and save them as named profiles for different document families.

08

An API integrators can build on

Upload, poll, read and edit findings, apply, download: the same REST API the interface uses, with an API key, documented in OpenAPI and served offline.

SECURITY & DEPLOYMENT

Your documents never leave the building

GDA is delivered as a signed, offline installation bundle for your own server. Every component that touches a document, from text extraction to the language model, runs inside that installation.

Everything runs on your host: text extraction, OCR, the entity model and the language model. After the one-time model download, the installation needs no internet connection.
Documents, extracted text and redacted copies are encrypted at rest with AES-256-GCM and deleted after the retention period you set, 30 days by default.
Named users with admin and reviewer roles. Every upload, decision, download and settings change is in the audit trail, which holds no document text and no personal data.
Signed installation bundle, hash-locked dependencies, pinned models, hardened containers, SBOM and vulnerability scan report with every release, and a security pack with checks your own team can run.
Runs on a CPU-only Linux server, an NVIDIA GPU server or Apple Silicon.
USE CASES

Where documents must be shared, but people must not be

Universities and research institutes

Real administrative, legal or financial records used for research or teaching, anonymised once with a documented review and reusable by the whole team.

Courts, lawyers and notaries

Decisions, deeds and case files shared for publication, research or training, with parties and witnesses pseudonymised consistently.

HR, payroll and procurement files

Contracts, payslips and offers handed to auditors, consultants or bidders without exposing employees and representatives.

Documents for AI and analytics

Build training sets, retrieval corpora and evaluation sets from real documents, and feed GRAG or GStT output onward, without putting people into a model.

Access-to-documents requests

Release the document, not the third parties in it.

Banks, insurers and telecoms

Loan files, claims and customer correspondence passed between departments and vendors under data minimisation.

LICENSING

One license per installation, nothing per user or per document

GDA is licensed per installation for a fixed term, with unlimited users and documents. We install it on your server from the signed bundle, or your team does with the included installer and runbooks. Upgrades, backups and restores are scripted.

FAQ

Answers before you ask

Does any document, or any part of one, leave our network?

No. Text extraction, OCR, the entity model and the language model run inside the installation on your server, on internal networks with no route out. There is no telemetry, no update check and no licence call home. The security pack explains how your own team can verify each of these.

Which formats does GDA handle?

Plain text, PDF with a text layer or scanned, and Word DOCX. The redacted copy comes back in the same format. Images and legacy .doc files are not accepted in this version.

Which personal data does it detect?

Names, home addresses, birth dates, ΑΦΜ, ΑΜΚΑ, IBAN, card numbers, Greek phone numbers, e-mail addresses, identity card and passport numbers, and optionally organisations, places and company tax numbers. Your own term lists add anything specific to you.

What hardware do we need?

One server of your own, physical or virtual, running Linux. Ordinary processors are enough; a graphics card makes the verification step faster. No cloud account and no internet connection are needed, and we size the machine with you during setup.

Can it run without the language model?

Yes. Verification can be switched off per profile. Detection then relies on patterns, the entity model and heuristics, runs faster, and errs towards redacting more. On the measured sets the language model's main effect is fewer false alarms, not more findings.

Who can use it, and what is logged?

Named users with two roles: administrators manage users, profiles and settings; reviewers upload, review, apply and download. Every action is in the audit trail with the acting user. Application logs never contain document text, personal data or client addresses.

Is pseudonymisation enough for the GDPR?

Pseudonymised data remains personal data under Article 4(5) of the GDPR. GDA is a measure of data protection by design and by default (Article 25) and of data minimisation, not an exit from the Regulation. Whether a given release is lawful remains your assessment; GDA gives you a consistent, reviewed and auditable way to make it.

See GDA on your own documents

Book a demo and we walk you through detection, review and output on sample documents. If it fits, we set up a 30-day evaluation on your own server.

Request a demo