Logo XGen-AI Smart Documents SL
Data protection

Protecting your data isn't a layer we add afterward.

It's part of how MIKA is built. This page explains, with the same precision we hold the rest of the site to, what we protect today and what we're still finishing documenting.

What data can MIKA process

The information you work with, as it is.

This is what MIKA can process. The more sensitive the data, the more it matters how it's protected before it reaches any AI model.

Unstructured documents

Invoices, contracts, forms and reports — uploaded in batch or connected from your usual sources, like Drive or SharePoint.

Handwritten documents

Clinical histories, expense notes and other handwritten documents, via handwritten OCR.

Data in connected databases

Oracle, SQL Server, MySQL or PostgreSQL — what they contain depends on your organization, not on MIKA.

Personal identifiers

Names, ID numbers, emails, phone numbers, banking data, health data and geolocation — the type of data MIKA actively detects and protects before any AI processing.

Images and audio

Medical images or audio with multiple speakers, when the use case requires it.

From entry to exit

How a document moves through MIKA.

These are the confirmed stages of processing. The last one doesn't have a defined public policy yet — we'd rather say so than make one up.

  1. Entry

    Batch upload, connection to your document sources or your databases.

  2. Processing

    Automatic pseudonymization, before the document reaches any AI model.

  3. Extraction

    OCR and extraction of structured fields.

  4. Storage

    On owned servers — today XGen-AI's own, with the direction of moving toward dedicated infrastructure per client.

  5. Use

    Document chat, semantic search, database queries, export.

  6. Retention and deletion

    Information pending publication

How we protect data

Before any AI model sees a document, it's already gone through this.

Pseudonymization
MIKA detects and tokenizes names, ID numbers, emails, phone numbers, banking data, health data and geolocation — before the document reaches any AI model, ours or a third party's. Only an authorized user can reverse that process.
Techniques applied
Masking, generalization, suppression and hashing, depending on the type of data and the level of protection it needs.
AI model isolation
The model doesn't access your complete database directly — it operates only on the already-treated fragments it needs to respond.
Traceability
The admin dashboard logs usage by company, user and level, with filterable, paginated logs.
Who can access it

Access organized across three levels.

MIKA's admin dashboard manages access through three distinct roles.

Superadmin

Global view of companies, users, products, licenses and logs across the whole system.

Company

Management and metrics for its own organization, with no access to other companies.

User

Individual panel, within the permissions their organization assigns them.

AI and privacy

No AI model sees a document as it arrived.

MIKA can use different language models — GPT, Gemini, Claude, or a model built in-house — depending on what each case needs.

Before any fragment goes out to any of those models, different layers of MIKA anonymize and hash it. What the model receives — ours or a third party's — is no longer identifiable personal data.

This applies every time, regardless of which model handles a given query.

Regulatory frameworks

Designed with the European regulatory framework in mind.

MIKA is designed to facilitate compliance with these regulations — this describes a design orientation, not a certification we hold.

Designed to facilitate compliance with
GDPREU AI ActNIS2DORALOPDGDD
Infrastructure aligned with
ISO 27001SOC 2 Type IIENS

We don't display certifications or seals because, today, we don't have a formal external audit to back them up. We'd rather say that clearly than suggest something we can't yet demonstrate.

Have specific questions about how MIKA handles your data?

Our team can answer with the same level of detail as this page — even where we're still finishing documenting something.