A user in Jakarta sent us a photo of German cookies she'd bought at a specialty shop. The ingredient list said "Schweineschmalz"—lard, from pigs. She had no way to know. Neither did our app, at first.

World

That was the day we realized a halal app that only works in English isn't really a halal app. It's a halal app for a privileged subset of the global Muslim community.

The Scale of the Problem

There are 1.9 billion Muslims worldwide. They live in countries where products are labeled in Arabic, Indonesian, Turkish, Urdu, French, Malay, and dozens of other languages. International food trade means that a German product might end up on shelves in Malaysia, and a Japanese snack might be sold in Saudi Arabia.

A woman in Morocco shouldn't need to learn German to know that "Gelatine" is gelatin. A man in Pakistan shouldn't need to recognize that "豚肉エキス" is Japanese for "pork extract."

We built a translation system to handle this. It took a year.

Why Real-Time Translation Doesn't Work

Our first approach was obvious: detect the language, translate to English, then verify. Google Translate exists. Problem solved, right?

No. Food ingredient terms are specialized vocabulary that general translators handle poorly. "Natrium benzoat" isn't natural language—it's a technical term that needs exact mapping to "sodium benzoate," not a creative interpretation. And translations need to be instant, not the 500ms+ that API calls require.

Worse, some ingredient names look the same across languages but mean different things in different regulatory contexts. "Gelatin" spelled identically in English, German, and Indonesian—but the sourcing norms differ by region.

The Translation Database

We built a pre-computed translation matrix. Every one of our 87,000+ base ingredients has entries in all 59 supported languages, creating millions of ingredient-language pairs. When OCR extracts "gélatine" from a French label, we look it up directly—no translation API, no interpretation, just a hash table lookup.

Language Family Languages Special Challenges
Latin Script English, French, German, Indonesian... Diacritics (é vs e)
Arabic Script Arabic, Urdu, Farsi Right-to-left, short vowels
Cyrillic Russian, Ukrainian Similar-looking characters
CJK Chinese, Japanese, Korean No word boundaries

Lookup time: under 10 milliseconds, regardless of language.

The Hard Cases

Some languages fight back.

Arabic and Hebrew — Right-to-left text that sometimes mixes in left-to-right ingredient codes (E471). ML Kit extracts these in visual order, which we have to reconstruct into logical order.

Chinese and Japanese — No spaces between words. We can't just split on whitespace; we need dictionary-based segmentation to find word boundaries.

German — Compound words that can be arbitrarily long. "Schweineschmalzersatz" (pork lard substitute) is one word. We have to recognize it as a compound and break it down.

Growing the Database

Translation coverage isn't a one-time project. New ingredients appear constantly, regional names vary, and OCR errors teach us about common misspellings we should handle.

Sources we draw from:

• Official food regulatory databases (FDA, EFSA, JAKIM)
• Academic food science literature
• User-submitted corrections
• Native speaker verification

Every unmatched ingredient in our logs is a candidate for addition. If users are scanning products with "natriummietabisulfiet" (Afrikaans for sodium metabisulfite) and we're failing to match it, that tells us we need Afrikaans coverage.

A halal app isn't global until it works in the languages global Muslims actually speak.