I watched a shared inbox tool go from 10 out of 10 to zero on ChatGPT just by rewording the question
I have been running a small experiment for a few weeks, and this one genuinely stopped me for a second. More and more buyers ask ChatGPT or Claude for a tool recommendation before they ever open Google. So I have been m
I have been running a small experiment for a few weeks, and this one genuinely stopped me for a second.
More and more buyers ask ChatGPT or Claude for a tool recommendation before they ever open Google. So I have been measuring what the models actually say when someone asks for the best tool in a category, and how much the answer moves when you change nothing but the wording.
Shared inbox software gave me the clearest example so far.
I asked two versions of the same question, 10 times each, on ChatGPT, Claude, Perplexity and Gemini. First the plain category label, "best shared inbox software". Then the way a real person actually types it, "how do we collaborate on shared email inboxes without losing our normal email workflow".
On the label question, Missive got named on all 10 runs. On the reworded one, ChatGPT named Missive zero times out of ten, while it kept naming Front and Help Scout on all ten. Same company, same site, same ten runs. The only thing that changed was the sentence.
To be fair it was not every model. Claude actually leaned the other way and named Missive more on the reworded question. But the ChatGPT drop was clean and it repeated across all ten runs.
The thing I keep taking from these is that a single check will lie to you. If I had asked once, seen Missive missing, and stopped there, I would have sworn they had some deep AI problem. They do not, they are 10 out of 10 on the other phrasing. One question on one model on one run tells you almost nothing.
And the tools that held did not do anything clever. Their pages just describe the actual job in the words a buyer uses, so the model can place them no matter how the question is asked.
I wrote up the full data, every number and every model, here: https://www.bersyn.com/blog/shared-inbox-phrasing-flip
I am running one of these teardowns roughly every week, poking at a different category. If that is your kind of rabbit hole, follow along. And I am curious, in your own category, is the recommendation this sensitive to phrasing, or did shared inbox just happen to be a swingy one?
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes â full credit and traffic to the original publisher.