The Tedious 90% of Localization QA

Localization QA is two jobs wearing one name. AI agents are better than humans at the repetitive 90%. Here is why, and what the other 10% still needs people for.

Cover Image for The Tedious 90% of Localization QA

Localization QA is two different jobs wearing one name.

The first job is tedious. Run the same forty user journeys in fifteen locales, on every release. Check that dates, numbers, currency, and plurals follow each locale's rules. Verify the glossary was respected, including the terms that are supposed to stay in English. Hunt for untranslated strings and raw i18n keys that leaked to production. Catch the button that truncates in German and the layout that overflows in Finnish. Confirm the RTL build did not just mirror the design but kept the flow usable.

The second job needs judgment. Does this headline land in Japanese, or does it read like a contract? Is the tone right for the Brazilian market? Would a native speaker wince?

The industry's standard answer to both jobs has been the same: a crowd of human testers. And for the first job, that answer has always been a bad fit, because the first job punishes exactly the things humans are worst at.

Why humans lose the tedious 90%

Not because testers lack skill. Because the work is hostile to human attention:

  • Fatigue. The fortieth checkout flow of the day gets less scrutiny than the first. Locale formats are precisely the kind of detail that fatigue eats first.
  • Sampling. Crowd cycles are expensive, so vendors sample: some locales, some flows, this quarter. What was not sampled is where the bug is.
  • Variance. Coverage depends on who picked up the job. Two testers flag different things on the same screen, and neither report is re-runnable.
  • Latency. A full crowd cycle takes weeks. Your release cadence is measured in days. So localization QA quietly falls out of the release process and becomes an occasional event.

An AI agent has the opposite profile. It does not tire on the fortieth flow. It applies each locale's formatting rules the same way every run. It follows the glossary consistently, including the harder discipline of not flagging the terms that are intentionally untranslated. And it covers every locale you ship, not a sample, because marginal coverage costs compute rather than tester-hours.

The numbers make the shape of this concrete. Our engineers typically author a localization regression suite in about two days. After that, it runs in about four hours, on every release. A typical crowd-testing cycle for the same checks takes about five weeks. Weeks become hours, sampled locales become all of them, and portal tickets become replayable Playwright tests.

The 10% we do not pretend to automate

Honesty matters more in QA than in most products, so here is the boundary as we see it.

AI is not the right tool for native-speaker taste: whether copy feels natural, whether humor survives translation, whether the tone fits the market. It is not the right tool for testing real payment instruments in each country. And it does not replace a physical device lab.

That is why Donobu's localization QA is a managed service rather than a tool we hand you. AI agents do about 90% of the work: your real user journeys, in every locale, on every release. Where AI is not enough, forward-deployed test engineers take over. And test engineers review every result before you see it, so you never receive unreviewed AI output. We automate the tasks, not the people.

What the findings look like

A finding is evidence, not an opinion: a screenshot of the failure state, the exact locale and step, and a replayable Playwright test. Your team can re-run the test after the fix and watch it go green, instead of debating a tester's screenshot from three weeks ago.

See it on your own site

If you ship in more than one language, start with the two free ways to see this in action:

  • Run the free localization check on any page you ship. It reads the declared language and hreflang tags, then flags untranslated text, mixed-language content, and locale-wrong formats.
  • Or go deeper: talk to us and get an LQA snapshot of your site across the locales you care about.

The tedious 90% is the part of localization QA that was never a good job for humans. We built the system that takes it off their plate. Read more on the Localization QA service page.