Dev.to Security 🔐 Cybersecurity 👁 0 📖 4 min read

How to check a Word document for black-bar redaction that is not really redacted

At the end of September a Nebraska TV station reported that a data center operator's state water and electricity filing had its figures "redacted" with a box that copy and paste undid: select the box, paste it into anoth

At the end of September a Nebraska TV station reported that a data center operator's state water and electricity filing had its figures "redacted" with a box that copy and paste undid: select the box, paste it into another document, and the numbers appear. Word files fail in a similar way, and in a few more.

A common way to "redact" a Word document is to select the sensitive text and give it a black highlight. It looks right on screen and in a PDF export. The text is still in the file.

A .docx is a zip archive, and the body lives in word/document.xml. Highlighting does not remove text from that file; it only adds formatting next to it. Anyone can select the black bar, copy it and paste it into a plain-text editor, or change the font colour and read it. Four things routinely survive when people think they have redacted a Word file:

  • Black on black. A run with <w:highlight w:val="black"/> or run shading with w:fill="000000", plus a black font colour, shows as a bar. The <w:t> element with the words is still there. (With the font colour left on automatic, Word turns the text white on a dark background, so the text stays readable: only the black-on-black combination looks redacted.)
  • Tracked deletions. Text deleted with Track Changes is wrapped in <w:del> with a <w:delText> element. It disappears from the page view but stays in the file until someone accepts the change.
  • Hidden text. A run with <w:vanish/> in its properties (Font > Hidden) is not shown or printed but is in the XML.
  • A black rectangle drawn over the text. The shape is a separate drawing element anchored to the paragraph; the text under it is untouched.

A small check

This script finds the first three. It reads only word/document.xml, so it ignores headers, footers, footnotes and comments, and it does not look at white-on-white text, paragraph or table-cell shading, or shapes. Treat it as a way to see the problem, not as a release gate. I wrote this article with AI assistance; the output shown below is from running the script on the test files described.

import html, re, sys, zipfile

def text_of(xml):
    return html.unescape("".join(re.findall(r"<w:(?:t|delText)(?: [^>]*)?>(.*?)</w:(?:t|delText)>", xml, re.S)))

def runs(xml):
    # every <w:r>...</w:r> with its properties and visible/deleted text
    for m in re.finditer(r"<w:r[ >].*?</w:r>", xml, re.S):
        r = m.group(0)
        props = re.search(r"<w:rPr>(.*?)</w:rPr>", r, re.S)
        yield (props.group(1) if props else ""), text_of(r)

def check(path):
    xml = zipfile.ZipFile(path).read("word/document.xml").decode("utf8")
    found = []
    for props, text in runs(xml):
        if not text.strip():
            continue
        black_bg = 'w:highlight w:val="black"' in props or re.search(r'<w:shd [^>]*w:fill="000000"', props)
        black_fg = re.search(r'<w:color w:val="000000"', props)
        if black_bg and black_fg:
            found.append(("black on black", text))
        elif "<w:vanish" in props:
            found.append(("hidden text", text))
    for m in re.finditer(r"<w:del [^>]*>(.*?)</w:del>", xml, re.S):
        found.append(("tracked deletion", text_of(m.group(1))))
    return found

if __name__ == "__main__":
    for kind, text in check(sys.argv[1]):
        print(f"{kind}: {text!r}")

Run it as python docx_redaction_check.py report.docx. On a test file with one of each failure it prints:

black on black: 'SECRET-ALPHA'
black on black: 'SECRET-BETA'
hidden text: 'SECRET-HIDDEN'
tracked deletion: 'SECRET-DEL'

and on a file with only ordinary text and a yellow highlight it prints nothing.

Fixing it properly

Formatting can be undone, so the fix is to remove the text itself, not to cover it:

  1. Delete the sensitive words (or replace them with a neutral marker such as [REDACTED]) while Track Changes is off, or accept or reject every tracked change first.
  2. Remove hidden text, and check headers, footers, footnotes and comments too.
  3. Save a new copy and run a check on that copy, not on the original. For anything that must stay confidential, also read the file's text in a plain-text view before you send it.

For a PDF that has to carry the redaction, export from the cleaned document and then check the PDF the same way: copy-paste from it.

Where this check stops

Colours that come from a style or a theme, rather than being set on the run, are invisible to the script above. Text colour and background can also differ by a hair (for example 010101 on black), which an exact string match misses. A real checker compares contrast, resolves styles, and reads every part of the package. The browser version below also flags automatic font colour on a black highlight, because Word draws that text white while other programs draw it black; the script does not.

I put a browser version of this idea, with those extra cases and a fixed-copy download, at https://hexloomlabs.com/redaction-check/?ref=devto-redaction. It runs in the page and uploads nothing; after it checks your files it reports how many network requests it made, and you can confirm that in your browser's Network tab.

📰 Read the original article on Dev.to Security

Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.