Two of our guide pages are instructions and one is a list, so only one of them emits a HowTo
Try this: curl -s https://pub-trivia.app/guides/how-to-host-a-pub-quiz \ | grep -o '"@type":"HowTo"' # "@type":"HowTo" curl -s https://pub-trivia.app/guides/pub-quiz-round-ideas \ | grep -o '"@type":"HowTo"' # no
Try this:
curl -s https://pub-trivia.app/guides/how-to-host-a-pub-quiz \
| grep -o '"@type":"HowTo"'
# "@type":"HowTo"
curl -s https://pub-trivia.app/guides/pub-quiz-round-ideas \
| grep -o '"@type":"HowTo"'
# nothing
Two pages in the same cluster, built from the same component, with the same metadata pipeline behind them. One of them is a numbered procedure a reader follows from start to finish. The other is a list of ideas you pick from. Only the first one is a HowTo, and the restraint is the part of structured data that is actually hard.
We run a 72 page marketing site for quiz software. There are six JSON-LD builders behind it, and the rule they all follow is one sentence long.
Every builder is fed the data the page renders
Structured data has one rule that matters more than the schema itself: it must
describe what is visibly on the page. Every builder here takes the same data
the page renders from, so the two cannot disagree: a FAQ block and its
`FAQPage` node are the same array, and a breadcrumb trail and its
`BreadcrumbList` are the same trail.
That is the comment at the top of the file, and it is an architectural claim rather than a style note. The usual way structured data goes wrong is not a malformed node, it is a correct node describing a page that has changed. Somebody adds three questions to a FAQ and the FAQPage still lists the original twelve, and now you have markup that is lying to a crawler on a page where every visible word is true.
Taking the same array the component maps over removes the possibility:
export function faqLd(entries: readonly FaqEntry[]) {
return {
'@context': 'https://schema.org',
'@type': 'FAQPage',
mainEntity: entries.map((entry) => ({
'@type': 'Question',
name: entry.question,
acceptedAnswer: { '@type': 'Answer', text: entry.answer },
})),
}
}
There is nothing clever in that function, and that is the whole point of it. It cannot drift because it holds nothing. Our FAQ page currently emits sixteen questions, which you can count from either end:
curl -s https://pub-trivia.app/faq | grep -o '"@type":"Question"' | wc -l
# 16
A one-item breadcrumb is not a breadcrumb
export function breadcrumbLd(path: string) {
const trail = breadcrumbFor(path)
if (trail.length < 2) return null
// ...
}
The homepage has no trail, and the temptation is to emit a BreadcrumbList with one item in it saying "Home". That is valid markup describing nothing, and it is the sort of thing that accumulates: a dozen nodes that are technically present and say nothing, on a site where you then cannot tell which markup is load bearing.
Returning null works because the component that renders these treats null as nothing to do:
export function JsonLd({ data }: { data: unknown }) {
if (!data) return null
return (
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(data) }}
/>
)
}
One place serialises a node into a script tag, which was about to be typed out on every content page. dangerouslySetInnerHTML is the documented approach in Next, and it is safe here for the reason it is always safe: everything passed in comes from the registry and the page's own module scope, and none of it comes from a request. The day any of this takes a URL parameter is the day that line needs a different answer.
HowTo only where the page is a procedure
/**
* An ordered set of steps, for the guides that genuinely are instructions.
*
* Only used where the page really is a numbered procedure a reader follows.
* "Pub quiz round ideas" is a list, not a HowTo, and marking it up as one would
* be describing the page as something it is not.
*/
export function howToLd({ node, steps, totalTime }: { /* ... */ })
This is the builder most likely to be abused, because HowTo is the one with a visible payoff. A list of ideas can be bent into steps if you are willing to call "pick a music round" a step, and the schema validator will not stop you.
The reason not to is not that Google will catch you, although it might. It is that the moment you start deciding markup by what you want rather than by what the page is, you have no principle left, and the next decision is worse. We publish fourteen guides under one hub. Exactly one of them emits a HowTo, with eight steps and totalTime: 'PT2H', because a quiz night takes about two hours and that page really is a procedure, from picking a night through to handing out the prize.
The totalTime is optional in the builder, and it is spread conditionally:
...(totalTime ? { totalTime } : {}),
Which is the same pattern as omitting a field entirely rather than emitting totalTime: null. An absent field means "we are not claiming this". A null field means "we are claiming this and the claim is nothing".
The author is the Organization, because we do not have a human to name
/**
* The author is the Organization rather than a person. A named human author is
* the stronger signal and we do not have one to name: inventing "Sarah, Head of
* Quizzes" to satisfy a schema field is exactly the kind of thing structured
* data is meant to stop. If a real person starts putting their name to these,
* this is where it goes.
*/
A Person author is a better signal than an Organization author. Every piece of SEO advice will tell you so, and the advice is correct. It is also an invitation to make somebody up, and plenty of sites have taken it, complete with a generated headshot.
So the author is an @id reference to our Organization node:
author: { '@id': organizationIdFor(SELF_APP) },
publisher: { '@id': organizationIdFor(SELF_APP) },
A reference rather than an inline object, pointing at the #organization node the site emits on every page, which is its own piece of machinery shared across four of our sites. Referencing by id means the author and the publisher are the same entity rather than two objects that happen to have matching fields, which is the thing JSON-LD's @id is for and the thing most implementations skip.
The field is a placeholder for a real answer rather than a permanent position. If somebody starts signing these guides, the line changes. Until then, the honest markup is the weaker one.
datePublished and dateModified are the same date
/**
* `datePublished` and `dateModified` are both the node's `updated` date. These
* are evergreen pages revised in place rather than dated posts, so the date a
* guide was first written is not a fact about it that anyone needs; the date it
* was last checked is.
*/
const iso = `${node.updated}T00:00:00Z`
This one gets argued about, so here is the reasoning. A guide to hosting a pub quiz is not a news item. Nobody benefits from knowing it was first drafted in March, and publishing an old datePublished on a page revised last week invites a crawler to treat it as stale content. Publishing a fake recent datePublished is worse.
Collapsing both onto the one date we actually track means there is one fact to keep true, and updated is bumped when the words change rather than when the file is reformatted. That same field is what the sitemap publishes as lastmod, so a page cannot claim one date to a crawler through the sitemap and another through its markup.
The offer is quoted in GBP even when the page is not
offers: offers.map((offer) => ({
'@type': 'Offer',
name: offer.name,
price: offer.priceGbp,
priceCurrency: 'GBP',
category: 'subscription',
url: absoluteUrl('/pricing'),
})),
Our pricing page converts for display, so a visitor in Germany sees euros. The structured data does not: it says £30.00 and £100.00 in GBP, always.
A per-visitor price in markup means publishing a different price on every crawl, which is a worse answer than a fixed one in the currency the plans are actually defined in. The conversion is a display convenience and the GBP figure is the contract, so the markup carries the contract. You can see both at once:
curl -s https://pub-trivia.app/features | grep -o '"priceCurrency":"GBP"'
# "priceCurrency":"GBP"
Then open the pricing page through a VPN exit somewhere in the eurozone and watch the visible number change while that line does not.
Six builders, and the ones we do not have
BreadcrumbList, FAQPage, SoftwareApplication, ItemList, Article, HowTo. That is the full set, and the absences are deliberate: no Review or AggregateRating, because we have no reviews and self-issued rating markup is the single most abused type in the vocabulary. No Product either, since SoftwareApplication already describes it and emitting both would be two nodes competing to be the same thing.
The general principle, if there is one: structured data is a set of claims about a page, and the useful discipline is to go looking for claims to delete rather than types to add.
Check it yourself
Every node on the site is in the HTML, so you can read all of it with curl and a JSON parser. The three pages worth comparing:
- pub-trivia.app/guides/how-to-host-a-pub-quiz emits Organization, Article, BreadcrumbList and HowTo.
- pub-trivia.app/guides/pub-quiz-round-ideas emits the first three and no HowTo, which is the decision this post is named after.
- pub-trivia.app/faq emits a FAQPage whose question count you can check against the questions you can count with your eyes.
If you want to check the claim in the title across the whole cluster rather than on the two pages I picked, this walks every guide in the sitemap and counts its HowTo nodes:
for u in $(curl -s https://pub-trivia.app/sitemap.xml \
| grep -o 'https://pub-trivia.app/guides[^<]*'); do
printf '%s %s\n' "$(curl -s "$u" | grep -c '"@type":"HowTo"')" "$u"
done
Fifteen URLs, one 1, fourteen 0s. And if you want to see what the pages are selling, the free tier needs no card.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.