Dev.to WebDev 🛠 Dev 👁 0 📖 8 min read

A Citation Gap Report in 70 Lines of Node.js, Using Real Directory Data

A citation gap report answers three questions: what share of the directories that matter is a business already on, which missing ones are worth the most, and where do the existing listings disagree about the business's n

A citation gap report answers three questions: what share of the directories that matter is a business already on, which missing ones are worth the most, and where do the existing listings disagree about the business's name, address or phone? You can build it in about 70 lines of dependency-free Node.js (20+, ES modules) from two inputs: a CSV of existing listings and a JSON target set. The only part that takes thought is the comparison. Strict string equality drowns you in false alarms, so the script normalizes before it compares.

The code below is what I ran, and the output is pasted from that run.

The inputs

The target set is real. It comes from the building and construction set in our catalog, which has 95 directories; only 32 of them have a domain rating (DR). Our public list of citation sites for builders shows the top 15 by relevance, so this set is a re-sort and not a copy of that page: I took the 12 with the highest DR, then added Bing Places explicitly as one of two global platforms every business should have. Foursquare is the other, and it already ranked in the top 12 at DR 91. Bing Places has no DR in the data, so it is stored as null, and the script has to cope with that. A DR tie at the cut line (Callupcontact and Checkatrade, both 74) was broken alphabetically, which is why Callupcontact is in. That tie-break is arbitrary, and for a Manchester builder Checkatrade, a UK trade directory, would be the more useful of the two. Swap it in if you reuse the set.

The business is fictional: "Example Builders Ltd", with a made-up Manchester address and a made-up phone number. targets.json carries it as a business object next to the directories array:

{
  "industry": "building-construction",
  "business": { "name": "Example Builders Ltd", "phone": "0161 496 0123",
                "address": "14 Example Road, Manchester M1 1AA", "fictional": true },
  "directories": [
    { "name": "Nextdoor", "url": "https://www.nextdoor.com", "domain_rating": 92 },
    { "name": "Foursquare", "url": "https://foursquare.com/", "domain_rating": 91 },
    { "name": "Bing Places", "url": "https://www.bing.com/maps/", "domain_rating": null }
  ]
}

That is an excerpt; the full file has 13 directories. The listings file is a CSV with deliberately messy but mostly harmless differences, because that is what real exports look like:

directory,url,name,phone,address
Houzz,https://www.houzz.com/professionals/example-builders,Example Builders Ltd,0161 496 0123,"14 Example Rd, Manchester M1 1AA"
Angi,https://www.angi.com/companylist/us/example-builders.htm,Example Builders Limited,+44 161 496 0123,"14 Example Road, Manchester M1 1AA"
Nextdoor,https://nextdoor.com/pages/example-builders,Example Builders Ltd,0161 496 0188,"14 Example Road, Manchester M1 1AA"
OpenStreetMap,https://www.openstreetmap.org/node/1000000001,Example Builders Manchester,0161 496 0123,"14 Example Road, Manchester M1 1AA"
Porch,https://porch.com/manchester/example-builders,Example Builders Ltd,0161 496 0123,"41 Example Road, Manchester M1 1AA"
Brownbook,https://www.brownbook.net/business/1000000002/example-builders,Example Builders Ltd.,(0161) 496-0123,"14 Example Rd., Manchester, M11AA"
Hotfrog,https://www.hotfrog.com/company/1000000003,Example Builders,0161 496 0123,"14 Example Road, Manchester M1 1AA"
Yell,https://www.yell.com/biz/example-builders-manchester-1000000004,Example Builders Ltd,0161 496 0123,"14 Example Road, Manchester M1 1AA"

Some of these differences are formatting (Rd for Road, +44 161 for 0161, Limited for Ltd) and should not count. Three are real: Nextdoor has a different phone number, Porch has the wrong house number, and OpenStreetMap has appended a place name.

The script

// Citation gap report: coverage, missing directories ranked by domain rating, NAP mismatches.
// Node 20+, ES module, no dependencies.
// Usage: node gap-report.mjs listed.csv targets.json
import { readFileSync } from 'node:fs';

// --- CSV: a minimal parser that understands quoted fields with commas ---
function parseCsv(text) {
  const rows = [];
  let row = [], field = '', quoted = false;
  for (let i = 0; i < text.length; i++) {
    const c = text[i];
    if (quoted) {
      if (c === '"' && text[i + 1] === '"') { field += '"'; i++; }
      else if (c === '"') quoted = false;
      else field += c;
    } else if (c === '"') quoted = true;
    else if (c === ',') { row.push(field); field = ''; }
    else if (c === '\n' || c === '\r') {
      if (c === '\r' && text[i + 1] === '\n') i++;
      row.push(field); field = '';
      if (row.some(v => v !== '')) rows.push(row);
      row = [];
    } else field += c;
  }
  if (field !== '' || row.length) { row.push(field); rows.push(row); }
  const [head, ...body] = rows;
  return body.map(r => Object.fromEntries(head.map((h, i) => [h.trim(), (r[i] ?? '').trim()])));
}

// --- normalizers: reduce a NAP field to the form we compare ---
const NAME_SUFFIX = /\s+(ltd|limited|llc|inc|co|company)$/;
const normName = s => s.toLowerCase().replace(/[^a-z0-9 ]+/g, ' ').replace(/\s+/g, ' ').trim().replace(NAME_SUFFIX, '');
const normPhone = s => s.replace(/\D/g, '').slice(-9); // last 9 digits: ignores +44, 0044, (0) and spacing
const ABBR = { rd: 'road', st: 'street', ave: 'avenue', ln: 'lane', dr: 'drive' };
const normAddress = s => s.toLowerCase().replace(/[^a-z0-9 ]+/g, ' ').split(/\s+/).filter(Boolean)
  .map(w => ABBR[w] ?? w).join('');   // joined without spaces so "M1 1AA" equals "M11AA"

const host = u => { try { return new URL(u).hostname.replace(/^www\./, ''); } catch { return ''; } };

// --- main ---
const [csvPath, targetsPath] = process.argv.slice(2);
if (!csvPath || !targetsPath) { console.error('Usage: node gap-report.mjs listed.csv targets.json'); process.exit(2); }
const listed = parseCsv(readFileSync(csvPath, 'utf8'));
const { business, directories } = JSON.parse(readFileSync(targetsPath, 'utf8'));

const findTarget = l => directories.find(d => host(d.url) === host(l.url) || normName(d.name) === normName(l.directory));
const covered = new Set(), outside = [], mismatches = [];
const truth = { name: normName(business.name), phone: normPhone(business.phone), address: normAddress(business.address) };

for (const l of listed) {
  const t = findTarget(l);
  if (!t) { outside.push(l.directory); continue; }
  covered.add(t);
  const shown = { name: normName(l.name), phone: normPhone(l.phone), address: normAddress(l.address) };
  for (const field of ['name', 'phone', 'address'])
    if (shown[field] !== truth[field]) mismatches.push({ directory: t.name, field, shown: l[field], expected: business[field] });
}

const missing = directories.filter(d => !covered.has(d))
  .sort((a, b) => (b.domain_rating ?? -1) - (a.domain_rating ?? -1) || a.name.localeCompare(b.name));
const pct = (100 * covered.size / directories.length).toFixed(1);

console.log(`Coverage: ${covered.size} of ${directories.length} target directories (${pct}%)`);
console.log('\nMissing, ranked by domain rating:');
missing.forEach((d, i) => console.log(`${String(i + 1).padStart(2)}. ${d.name.padEnd(16)} DR ${d.domain_rating ?? 'n/a'}`));
console.log('\nNAP mismatches:');
if (!mismatches.length) console.log('none');
for (const m of mismatches) console.log(`- ${m.directory}: ${m.field} shows "${m.shown}", expected "${m.expected}"`);
if (outside.length) console.log(`\nListed outside the target set: ${outside.join(', ')}`);

Three decisions are worth explaining.

Matching a listing to a target. findTarget compares hostnames with www. stripped, then falls back to the directory name. The fallback matters because a real export will not always agree with your target set about the domain: a listing on a country domain such as houzz.co.uk would never match a target stored as houzz.com.

Normalizing before comparing. Names drop punctuation and a trailing ltd, limited, llc, inc, co or company. Phones keep only digits and compare the last nine, which makes +44 161 496 0123, (0161) 496-0123 and 0161 496 0123 identical without a phone-number library. Addresses expand a few abbreviations and join without spaces, so M1 1AA equals M11AA.

Ranking what is missing. The sort puts the highest DR first and unrated sites last, then breaks ties by name. The ?? -1 makes that rule explicit: an unrated site sorts last, and a missing field cannot turn the subtraction into NaN and make the order unpredictable.

The output

Coverage: 7 of 13 target directories (53.8%)

Missing, ranked by domain rating:
 1. Foursquare       DR 91
 2. HomeAdvisor      DR 91
 3. CHECK24 Profis   DR 79
 4. Storeboard       DR 76
 5. Callupcontact    DR 74
 6. Bing Places      DR n/a

NAP mismatches:
- Nextdoor: phone shows "0161 496 0188", expected "0161 496 0123"
- OpenStreetMap: name shows "Example Builders Manchester", expected "Example Builders Ltd"
- Porch: address shows "41 Example Road, Manchester M1 1AA", expected "14 Example Road, Manchester M1 1AA"

Listed outside the target set: Yell

Coverage is 7 of 13, or 53.8%. The Yell row, a listing that is not in the target set, is reported rather than silently counted, so a spreadsheet that includes extra directories does not inflate the percentage.

The ranked list is the useful part. HomeAdvisor and Foursquare are the two highest-rated gaps, both DR 91, but HomeAdvisor is a North American directory, so for a Manchester builder it is the first row to cut, not to chase. That is the weakness of ranking by DR alone. Foursquare is worth closing first among the global platforms; the walkthrough in adding or claiming a Foursquare listing covers the search-before-you-add step that avoids creating a duplicate place. Bing Places sorts last only because the data has no rating for it. The report prints n/a rather than a zero for the missing rating. The notes on claiming Bing Places for Business explain the import-from-Google route and verification.

On this fixture, a strict string comparison of every field in the seven matched listings flags 10 mismatches. The normalized comparison flags 3. That gap is the argument for writing the normalizers: a report full of Rd versus Road gets ignored, and the real wrong phone number sits in the noise.

Limits you should know about

  • The last-nine-digits rule suits UK and US style numbers. Countries with shorter subscriber numbers need a real parser.
  • The abbreviation map is English and short. Extend it for your market.
  • Joining address words without spaces is a shortcut that can in principle equate two different addresses (1 4 and 14). It is good enough to screen, not to give a verdict.
  • Domain rating is a third-party snapshot, so treat the ranking as a prioritization aid, not a score of how much a listing will help.
  • The fixture is a raw slice of the highest-rated builder sites, which mixes country-specific directories (HomeAdvisor and Porch are North American, CHECK24 Profis is German) with global ones. Filter the target set by the business's country before you use a report like this for real work.

Where to take it

A --json flag and a non-zero exit code on mismatches would let a scheduled job fail loudly. Running the report across many clients is where the manual version gets tedious; the agency citation workflow describes how we run citation building for many client locations from one account, and the country-optimized lists and done-for-you building are covered on the features page.

I work on Citation Builder, an SEO citation builder that ranks citation sites by country and industry and builds the listings for the business owner. The data in this article comes from its directory catalog; the business is fictional.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.