Our curation prompt offered 'vegetable' and not 'onion', which is how an ingredient becomes recognised and unscoreable
Munchable reads the ingredient list off a food label and says whether the product suits the gut conditions you manage. Labels are printed in every language and every house style, so a lot of our work is an ingredient tax
Munchable reads the ingredient list off a food label and says whether the product suits the gut conditions you manage. Labels are printed in every language and every house style, so a lot of our work is an ingredient taxonomy: a graph of canonical ids, with aliases for the thousands of ways a pack can spell the same substance.
Unknown words from real labels go through a curation job. A model sees one slug at a time with its prevalence and some label context, and answers with one of four actions: alias it to something we already know, file it as a new entry under one or two parent concepts, mark it as noise, or decline. Then deterministic guards decide whether the answer is allowed to merge.
Two of those guards were broken in the same direction, and the direction matters more than the bugs. Both failures turned an unknown ingredient into a recognised one that no rule can ever reach.
Why a recognised ingredient can be worse than an unknown one
Rules key on ids and inherit downwards. If a rule set names a concept, everything beneath that concept in the graph is covered by it, which is the whole reason a graph is worth maintaining.
Naming an ingredient also does something else: it raises the engine's confidence in the row, because confidence is partly a count of words the engine could not place.
Put those together. File a new onion preparation under a general vegetable concept and you have made the word recognised, raised the confidence of every product carrying it, and placed it somewhere no rule reaches. The user gets a more confident answer that has quietly stopped considering an ingredient it would have considered yesterday. An unknown word is visible; this is not.
The list we handed the model named the general and skipped the specific
New entries need parents, and the prompt offers a list of concepts to choose from. The list was hand-written, and it looked sensible:
export const PARENT_CONCEPT_NAMES: readonly string[] = [
'cheese', 'dairy', 'milk', 'cream', 'butter', 'yogurt', 'fermented milk products',
'oil and fat', 'vegetable oil', 'animal fat',
'sweetener', 'sugar', 'syrup', 'honey', 'sugar alcohol', 'polyol',
'spice', 'herb', 'condiment', 'sauce', 'vinegar', 'salt',
'vegetable', 'root vegetable', 'legume', 'pulse', 'mushroom', 'seaweed',
'fruit', 'dried fruit', 'fruit juice', 'fruit puree', 'juice',
// ... 91 names in total
];
Ninety-one names, of which eighty-two resolve to a real id today. I measured how many of those eighty-two sit anywhere in the rule graph, meaning they are a rule key or inherit one:
general concepts offered as parents: 82
of those, reachable by any rule: 18
Read the list again with that in mind. It names vegetable and not onion. fruit and not apple. nut and not almond. No garlic, no tomato, no citrus. Those are exactly the words a gut-health rule set cares about, and a model told to prefer the list will dutifully pick the general term, because the specific one was never on offer.
So the defect was not in the model's answers. It was in the menu.
The fix is not a longer list, it is a derived list
Typing out the missing names would have fixed it for a week. The rule maps change; a hand-written mirror of them drifts, and this one would drift silently because nothing reads it back.
So the names come out of the rule maps themselves:
function ruleConceptNames(): string[] {
const out: string[] = [];
for (const id of [...HAND_RULE_KEYS].sort()) {
if (!isKnownTag(id)) continue;
const name = displayName(id);
if (/^E\d/.test(name)) continue; // ids, not concepts anyone names a parent
out.push(name);
}
return out;
}
export function parentConceptNames(): string[] {
if (!parentNames) {
const all = [...new Set([...PARENT_CONCEPT_NAMES, ...ruleConceptNames()])];
parentNames = all.filter((n) => resolveName(n) !== null);
}
return parentNames;
}
Today that offers 305 concepts instead of 82, and 223 of them arrived from the rule maps without anybody typing them. Add a rule tomorrow and the concept it keys on is on the menu the same night.
Two details in there are deliberate. E-numbers are dropped, because nobody names an E-number as a parent and the prompt already tells the model to prefer the number for additives. And the whole list is filtered through resolveName, so it can never suggest a parent that the guards would then reject: an option that cannot be accepted is a trap, not a choice.
The prompt stopped being neutral about generality at the same time:
Parents should be drawn from the concept list below whenever one fits, and
pick the MOST SPECIFIC one that is true: "onion" rather than "vegetable" for
an onion preparation, "almond" rather than "nut", "apple" rather than "fruit".
A parent that is too general loses what the ingredient actually is.
The guard that was a no-op exactly where the job runs
There was already a guard for this. It refused any decision that would move a health-relevant ingredient out of the rule graph, and it had been there for months.
It keyed on what we call the lexical resolution: what the slug's own words resolve to through our multilingual lexicon. Which is a strong signal, and it is null for precisely the words this job exists to process. A word the lexicon can already resolve is not in the backlog. The guard was running on every input and firing on none of them.
The fix is a fallback that scans the model's own proposed English name for a known ingredient, the same scan we use to catch a real substance hiding inside a boilerplate phrase:
const healthSignal = lexical ?? hiddenIngredient(normName(p.english_name), lang);
So "onion granules" now yields our onion id even though the slug itself resolved to nothing, and the guard has something to protect:
if (healthSignal && isHealthRelevant(healthSignal) && !parents.some(isHealthRelevant)) {
return reject(`would drop ${healthSignal} out of the rule graph`);
}
And on the alias path, the same signal with one extra case:
if (healthSignal && isHealthRelevant(healthSignal)) {
if (BLAND_TARGETS.has(to)) {
return `would hide a health-relevant ingredient (${healthSignal}) behind ${to}`;
}
if (!isHealthRelevant(to)) return `would drop ${healthSignal} out of the rule graph`;
}
Aliasing something to water or salt is how a trigger disappears most quietly, so the bland targets get their own message.
The half of the guard that checks lineage, meaning "does this claim contradict what the word already means", deliberately keeps using the lexical signal only. Checking the model's proposed name against the model's proposed name is circular, and it would reject honest compound names: a sunflower lecithin would be refused for not sitting under sunflower, when the additive id is exactly where it belongs.
What I take from both bugs
The options you hand a model are a schema, so derive them. A prompt that enumerates allowed values is a contract between the model and the system that will validate its output. Two hand-written halves of one contract will disagree, and the disagreement shows up as plausible answers that pass validation and do the wrong thing. We have made this mistake in two other places where a list derived from nothing drifted.
A guard keyed on your strongest signal may be structurally absent on the inputs that reach it. The lexical resolution was the right thing to check and the wrong thing to require, and the test suite was happy because every fixture had one. When you write a guard, ask what the inputs look like at the point it runs, not what they look like in the example you had in mind.
Both are specific cases of the thing that makes AI-assisted data curation survivable at all: the model proposes and the engine disposes, with guards that have the last word. A guard that cannot fire is not a guard, it is a comment. Relatedly, knowing a word is not understanding it is the first version of this same failure we shipped and fixed.
Go and check our homework
Every public answer page runs the production engine over our taxonomy, so each one is a live readout of whether a concept is reachable by a rule:
- Is onion low FODMAP?
- Does onion cause reflux?
- Is almond ok with gastroparesis?
- The whole set: munchable.app/answers
Then open app.munchable.app, pick a condition, and search for something with a compound ingredient list, a stock cube or a flavoured crisp. The reasons under the verdict name the ingredients that were considered. Anything filed in the wrong place in our graph would simply be missing from that list, which is why the menu the model chooses from gets derived rather than typed.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.