March Code
Translation and localization

Polyglot: an AI platform for a translation agency where the AI never sees personal data

How to use cloud AI models legally on documents with passports, court records and medical data: personal data is replaced with masks on the client's own server, only anonymized text goes to the cloud, and a second AI model checks every machine translation, while a human stays accountable for quality. Built for a translation agency in the EU: 9 stages from scan upload to delivery with an electronic signature, 24 languages, 1.5 million translated segments in translation memory.

Client: Translation agency, European Union (Latvia)AISaaSAutomation
Polyglot: an AI platform for a translation agency where the AI never sees personal data

0

personal data in the cloud: the AI sees only masks

4 of 209

segments the editor reads after the double AI check

1 minute

to create an order instead of assembling it from emails

2–3 months

from spec to production use

01

Challenge

Our client is a translation agency in the EU (Riga, Latvia). It works for government agencies and corporations: hundreds of orders a month, 24 languages, documents with special categories of personal data such as passports, court files and medical reports. The work is governed by ISO 17100, the European standard for translation services, and by GDPR. A file with someone's passport sent to the cloud as is means a risk of a fine of up to 4% of annual turnover and losing access to public contracts.

The agency's usual toolkit (email, Excel and the memoQ translation software) stops working at these volumes:

  • Project managers drown in routine. Every order is put together by hand from emails: files, rates, deadlines, finding a free translator.
  • Files with personal data go to contractors and cloud services unprotected. The regulator doesn't care that “everyone does it”: the agency is liable.
  • Prices are calculated by hand from discount tables, and mistakes go both ways: either the job runs at a loss, or the client disputes an inflated invoice.
  • “Who approved what” has to be pieced together from email threads. In a dispute over quality or deadlines, there's no evidence.

Simply “plugging in ChatGPT” wasn't an option. No serious client trusts machine translation without an independent check, and no lawyer will sign off on sending passport data to a cloud AI in plain form: all the liability would fall on the agency. The agency needed a system where AI speeds up every stage, a human is accountable for quality, and personal data never leaves the company's server.

Not in the EU? Swap GDPR for your local data protection law: for healthcare, legal services, finance and HR data, the problem is exactly the same. Below is how we solved it.

Screenshots are from a demo environment: projects, clients, names and documents are test data.

02

Solution

One pipeline instead of email, memoQ and Excel

Polyglot takes every document through nine stages, from scan upload to a package with an electronic signature. The system recognizes scans, splits the text into segments, fills in existing translations from the accumulated memory, sends new text to an AI model, has a second AI model check it, and a human approves the result. At every stage you can see where the document is, who owns it and what is waiting for a decision.

Three workspaces, three headaches gone

Project manager
All production on one screen
  • An order takes a minute to create: the system picks the translation memory, glossaries and translators for the language pair
  • A single decision queue grouped by project, so nothing gets lost in email
  • You see what's close to its deadline, what's overdue and what's waiting for the client
Translator
Fair job offers and a convenient editor
  • A job offer shows the rate, volume and deadline before you accept it: the first to accept gets the job
  • A workspace in the browser: translation memory, glossaries, suggestions, team comments
  • An earnings dashboard with an export for invoicing: the month's total adds itself up
In-house editor
Quality control under ISO 17100
  • The recognized scan is checked before translation starts, so garbage never enters the pipeline
  • A final delivery sign-off with explicit confirmation
  • The four-eyes principle is enforced by the server, not by trust: the system won't let you close a review of your own work

Anonymizing personal data: how AI translates other people's passports without seeing them

1. Masks instead of names, on the agency's server

The system finds names, addresses, passport and account numbers in the document (including local formats with checksums) and replaces them with masks before anything is sent to the cloud

2. Translation in the cloud

Only anonymized text goes out. The AI model translates it using the client's glossaries; the mask mapping table is encrypted and never leaves the server

3. Check by a second AI model

A second model from a different vendor checks the first one's translation and flags doubtful spots with a suggested fix. The editor works through the flags instead of rereading everything

4. Data restored, document delivered

Real values go back into the text on the agency's server. The document is rebuilt with its original formatting: DOCX, PDF and an EU-standard electronic signature package

What else sets the platform apart

  • Fair prices without manual math. Repetition discounts are calculated automatically: the more matches with the accumulated translation memory, the cheaper the order for the agency's client. The system simply won't issue a quote below the agreed minimum.
  • Years of translations weren't lost. The databases built over years of work were moved from memoQ and Trados in a single file: 1.5 million translated segments. Every delivered order adds to its own client's memory, and different clients' databases never mix.
  • AI under control, with a price tag. The cost of every AI call is known; legal text goes to one model, technical text to another. If the AI provider goes down, a backup model takes over on its own, and the cloud can be switched off with one button.
03

The system from the inside

Real screens of a working system, not mockups. Click to take a closer look

Production pipeline
The manager's dashboard shows all production and everything waiting for a human decision: what to review, whom to offer a job, what to deliver to the client and what's running late. No more status meetings
A new order in a minute: the system picks the client's translation memory and glossaries, detects the document type and calculates the price up front
Decision queue: everything awaiting approval, grouped by project, with how long each task has been waiting. A review can only be approved after opening the document, because the four-eyes principle is enforced by the server, not by trust
AI layer and data protection
Anonymization before the cloud: names, document numbers and addresses are found and masked on the agency's server. On the right is every detected piece of personal data; any display of a real value goes into the access log
AI translation and an independent check: the cloud sees masks instead of personal data. A second AI model checks the first one's translation and flags doubtful spots, here 4 flags across 209 segments
AI model catalog with prices and routing: legal text goes to one model, technical text to another. The cost of every AI call is tracked in euros
Data protection center: a catalog of personal data recognizers by type and language, a “cloud stores nothing” mode that turns on with one switch, and a full access log
Editor and translation memory
Translator workspace (CAT editor): translation memory, segment statuses, protection of tags from accidental edits, team comments. All in the browser, no memoQ install needed
Translation memory: 1.5 million translated segments in 24 languages. Every delivered order automatically adds to its client's database, and client databases never mix
Money and translators
The price is known before the order is created: discounts for translation memory matches, four ways to count volume, the translator's payout. The system won't issue a quote below the agreed minimum
Translator dashboard: job offers show the rate, volume and deadline before you accept. The first to accept gets the job; earnings and an invoicing export are in the personal account
04

Results

0

personal data in the cloud: the AI sees only masks

Names, addresses and document numbers are replaced with masks on the agency's server before any call to the AI model. Every time a real value is displayed, it's recorded in the access log

4 of 209

segments the editor reads after the double AI check

The first model translates; the second, from a different vendor, checks it and flags doubtful spots. The editor works through the flags instead of rereading the whole document

1 minute

to create an order instead of assembling it from emails

The system picks the memory, glossaries and translators for the language pair and calculates the price up front. 24 languages, 1.5 million translated segments, with databases moved from memoQ and Trados in a single file

2–3 months

from spec to production use

The core (pipeline, roles, translator workspace) went live in 2–3 months, with improvements every week since. Every release first passes 211 automated tests

05

In-depth breakdown

Why this case matters

  • AI the system is accountable for, not “the model's mood”. Two models check each other, every call has a known cost, and if a provider fails, a backup model takes over automatically. The client knows the cost of every order; the data and the process belong to the client, not to the AI vendor.
  • Sensitive data, taken seriously. Documents with special categories of personal data never leave the client's server in plain form. The architecture was designed from GDPR and ISO 17100 requirements up, not retrofitted to them after the fact.
  • A deep dive into the industry. We learned how the agency works down to the details, from industry discount grids and the four-eyes principle to government requirements for electronic signatures. That's how we approach any industry: process first, code second.
  • A custom system vs. subscriptions. Its own pricing rules, client data on its own hardware, independence from the AI vendor, and no per-seat fees that grow with every new hire.
  • Speed without cutting corners. The client saw the first working screens within a few weeks, and the core went live in 2–3 months. Every release runs 211 automated tests first, so updates don't break the agency's work.

The same architecture for any regulated data

Swap GDPR for your own data protection rules, and the problem will feel familiar to anyone who handles medical records, contracts, court documents or HR data. No personal data goes to the cloud, only masks, so there's no cross-border transfer of personal data and no processing of it by a third-party AI model. The setup was built for GDPR, where fines reach 4% of annual turnover. This approach opens up powerful large language models (LLMs) to fields where they used to be off the table: healthcare, law, finance, HR. This is exactly the kind of setup we design as part of our AI integration service, and the platform itself follows the playbook of custom SaaS development.

FAQ

How long does it take to build a platform for a translation agency?

The core (stage pipeline, roles, translator workspace) takes 2–3 months to go live. After that, the platform grows in one-week iterations with live releases: the client sees progress in working screens, not in reports.

Does the data go to OpenAI or Anthropic?

Only anonymized text: names, numbers and addresses are replaced with masks on the client's server before any call to the cloud. The mapping table is encrypted and never sent out. So there's no cross-border transfer of personal data and no processing of it by a third-party AI model, which takes care of the main data protection concern. There's also a mode with no AI at all and a “cloud stores nothing” switch.

How much does a platform like this cost?

The price depends on the number of seats, how deep the AI layer goes and the data protection requirements, so there's no honest one-size-fits-all price list. We work out a preliminary range for free on the first call: you walk us through your process and leave with an architecture sketch and a ballpark budget.

Who owns the code, and who supports the system after launch?

The source code, data and documentation belong to the client: under the contract we hand over the entire repository. Support is optional: our team can keep developing the system in iterations or hand it over to your developers.

Why not an off-the-shelf SaaS like memoQ, Phrase or Smartcat?

If your process is standard, a ready-made service is faster and cheaper, and we'll tell you so on the first call. A custom platform wins when data ownership, non-standard pricing rules, public-sector requirements and AI under your own control all matter. Here, all four did.

Need a system like this for your process?

We design and build industry platforms with an AI layer end to end, from process analysis to launch and support. Leave a request below: on a 30-minute call we'll show you a live Polyglot demo and sketch the architecture and a budget range for your task for free. If an off-the-shelf SaaS covers your process, we'll tell you so plainly.

Project technologies

Python 3.12FastAPISQLAlchemy 2PostgreSQL 16React 18TypeScriptTanStack QueryTailwindClaude / OpenAI APIMicrosoft PresidioSSEDocker ComposeCaddyWireGuardASiC-E (e-signature)

Want a similar result?

Pilot from $3,900, prototype free

Prices are indicative and not a binding offer.

We build an agent prototype for one scenario using your examples, so you can judge answer quality before signing.

Step 1 of 2 · Task