Dev.to WebDev 🛠 Dev 👁 0 📖 6 min read

Building a Chrome extension that reads Hanifi Rohingya webpages in Latin script

Rohingya has two writing systems in active use. Hanifi Rohingya is a right-to-left script that has been in Unicode since 2018 (block U+10D00–U+10D3F). Rohingyalish is Rohingya written in the Latin alphabet. Many Rohingya

Rohingya has two writing systems in active use. Hanifi Rohingya is a right-to-left script that has been in Unicode since 2018 (block U+10D00–U+10D3F). Rohingyalish is Rohingya written in the Latin alphabet. Many Rohingya speakers read Rohingyalish comfortably but struggle with Hanifi, and the reverse is true for others.

At RohingyaLanguage.org we already had a Hanifi ↔ Rohingyalish script converter. But copying text from a webpage into a converter and back gets tiring fast. So I built Rohingya Reader, a Chrome extension that converts Hanifi text in place, on the page you're reading.

This is transliteration (changing the writing system), not translation. The words stay Rohingya.

👉 Rohingya Reader is coming soon to the Chrome Web Store (it's in review). Until then, the same converter is available online.

Here's what turned out to be harder than expected, and how I handled it.

The requirements

  • Private by default. Everything runs on the device. No backend, no API, no analytics.
  • One converter. The extension imports the website's converter module directly. There's no copy of the mapping tables to drift out of sync.
  • Don't break the page. Links, event handlers, forms and the website's own JavaScript must keep working.
  • Reversible. "Show original" must put back exactly what was there.
  • Minimal permissions. No access to every website at install.

1. Detect Hanifi by code point, not by lang

Pages rarely mark Hanifi text with a correct lang attribute, and fonts vary. The only reliable signal is the characters themselves:

export const HANIFI_TEST = /[\u{10D00}-\u{10D3F}]/u;

Each text node is split into maximal runs of Hanifi code points. Only those runs go to the converter. English, Arabic, Bengali, emoji and punctuation are copied through unchanged. A property test checks, on thousands of random mixed strings, that the result is identical to running the website's converter on the whole string.

2. Change Text.data, never innerHTML

The fastest way to break a website is to rewrite innerHTML. That destroys event listeners, framework state and element identity. Rohingya Reader edits only the data of existing text nodes. No elements are created, moved or removed, and no attributes change.

For every node it changes, it remembers two strings: the original text and exactly what it wrote. That second value solves two problems:

if (rec && node.data === rec.written) return; // our own change: skip (no observer loop)
if (rec) records.delete(node);                // the *website* changed it: new source text

Restoring works the same way. A node is only put back if it still contains what the extension wrote. If a chat app or feed has replaced the text since, the website's version wins. Records sit in a WeakMap and a set of WeakRefs, so nodes the page removes can still be garbage-collected.

3. Dynamic pages, without jank

Feeds, comments and infinite scroll mean a one-time conversion isn't enough. A MutationObserver queues only the nodes that were added or changed. It never rescans the whole document. Work runs in slices of about 12 ms, so the page stays responsive. In an end-to-end test, a page with 5,000 paragraphs converted in about 0.4 seconds with zero long tasks.

Skipped entirely: script, style, inputs, textareas, contenteditable regions, role="textbox", code blocks, and common editors such as CodeMirror, Monaco and ProseMirror. Changing text under someone's cursor is not a feature.

4. Words split across elements

Hanifi uses combining signs for tone and length. Sites often wrap parts of a word in separate elements (a highlighted letter, or a tone mark in its own <span>):

𐴑𐴝<b>𐴤</b>

Converting each text node on its own gives a different result from converting the whole word, because the converter looks one character ahead. So adjacent inline text nodes are converted as a group. The node boundary may move by one or two code points, but only to a position where splitting provably doesn't change the converter's output. The converter itself serves as the check. If no safe split exists, the text stays in Hanifi rather than being guessed.

5. Right-to-left meets left-to-right

Hanifi is right-to-left; Rohingyalish is left-to-right. Drop Latin text into a dir="rtl" paragraph and the Unicode bidi algorithm will happily reorder the words around punctuation and links.

Setting dir="ltr" on elements would change the layout and touch attributes. Instead, the extension wraps each converted passage in invisible Unicode isolate characters, U+2066 (LRI) … U+2069 (PDI), written into the existing text nodes. One isolate covers the whole passage, even when it spans a link, so word order stays correct. The paragraph keeps its alignment, and dir="auto" elements fix themselves once their first strong character is Latin.

6. Permissions users can trust

The manifest requests activeTab, scripting and storage, plus optional_host_permissions. Nothing is granted at install:

  • Convert this page uses activeTab, which only works after you click the toolbar button.
  • Always convert this website asks Chrome for access to that exact hostname only (*://example.com/*), then registers a content script with chrome.scripting.registerContentScripts.

Three things can drift apart: the saved preference, the granted permission, and the registered script. Users can revoke access in Chrome's settings, and registrations can be lost when the extension updates. A background worker brings them back into agreement on install and update, at browser start-up, on permissions.onAdded and onRemoved, and whenever the popup opens.

One gotcha: Chrome's permission prompt can close the extension popup, and the popup's promise dies with it. The popup therefore tells the service worker it is about to ask, and the worker finishes the job in permissions.onAdded.

7. Testing a real extension

  • Unit tests (Vitest + jsdom): converter equivalence, restoration, website edits after conversion, split words, direction marks, and the permission lifecycle against a fake Chrome API.
  • End-to-end tests (Playwright): the real unpacked extension runs in Chromium. Branded Chrome 137+ no longer supports --load-extension, so Chromium it is. The tests cover right-to-left word order measured from actual layout, feeds, editable fields, axe accessibility scans of the popup, browser restarts, and revoking access through Chrome's own chrome://extensions toggle.

Automation can't click the toolbar button or answer Chrome's permission prompt. Those steps go into a short manual checklist rather than being claimed as tested.

Technical vs linguistic correctness

The tests prove the extension reproduces the converter exactly. They don't prove the converter matches how every writer spells Rohingyalish. The converter is rule-based and approximate:

  • Some tone signs are simplified.
  • Sakin has no Latin equivalent.
  • Some letter distinctions merge.

When it approximates, the popup says so and shows the original Hanifi next to the result.

A Rohingyalish → Hanifi → Rohingyalish round trip is a useful consistency check, not proof of accuracy. Across 9,200 distinct dictionary words, 187 don't round-trip, and all of them come from a few documented spelling conventions (ts, sh, and ng/ny versus ñg/ñy). Examples reviewed by fluent readers are the next step.

The dataset

The same converter and dictionary power an open dataset on Hugging Face:

Rohingya Hanifi–Rohingyalish–English Lexicon (CC BY 4.0)

  • 15,926 rows linking English headwords, Rohingyalish and Hanifi Rohingya
  • Built from 6,510 English dictionary entries
  • Every row passed an exact Rohingyalish → Hanifi → Rohingyalish round trip. Rows with converter warnings were excluded.
  • The Hanifi is converter-generated and not individually reviewed, so treat it as consistent, not as verified spelling.
from datasets import load_dataset

ds = load_dataset("rohingyalanguage/rohingya-hanifi-rohingyalish-english", split="train")
print(ds[0])

It's useful for transliteration experiments, search normalization, and low-resource NLP work on Rohingya.

Try it

If you read both Hanifi and Rohingyalish and want to help review conversions, or you run a website with Hanifi content and something doesn't look right, I'd love to hear from you: [email protected]

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.