Finding one
The methods disagreed about visibility in one out of six comparisons.
Our panel contained five selected businesses in every market context: a market leader, major challenger, value provider, premium or specialist provider, and smaller provider. For each business, we looked across eight unnamed buying questions and asked a simple question: was this business recommended at least once?
That produced 500 business-by-market comparisons. The API with web search and ChatGPT Search gave different answers in 82 of them: 16.4%, with a 95% confidence interval of 12.8% to 20.2%.
| Classification result | Comparisons | Rate |
|---|---|---|
| Both methods agreed | 418 | 83.6% |
| One recommended the business; the other did not | 82 | 16.4% |
We made the test stricter by requiring at least two appearances among the eight questions. The disagreement rate barely moved: 16.6%. The result was not simply created by isolated one-off mentions.
Finding two
Sometimes the apparent market winner changed too.
We counted how often each of the five sampled businesses was recommended across the eight questions in its market. In 96 contexts, both methods recommended at least one panel business. Their winning sets had no business in common in 14 of them, or 14.6%.
Some contexts had ties. When we isolated the 48 contexts where both methods produced one clear winner, they named different businesses in 11 cases, or 22.9%.
That number needs a sensitivity warning. When we required each winning business to appear on at least two of eight questions, the estimate fell to 16.3%. At a four-of-eight threshold, it fell to 6.9%. The lesson is not that 22.9% is a universal error rate. It is that the winner changed often enough to matter—and platforms should show how much evidence supports a ranking.
Three examples
What a changed winner looks like to a marketer.
A New York pest-control company went from invisible to the panel winner.
The API with web search recommended Control Exterminating on four of eight questions and OnGuard on none. ChatGPT Search recommended OnGuard on three and Control on two.
API with webAPI diagnosis: OnGuard has a fundamental discoverability problem and should benchmark Control.
ChatGPT SearchSearch diagnosis: OnGuard leads the sampled customer experience and should learn how to defend it.
“A business cannot sensibly be treated as both invisible and the market leader without explaining the measurement route.”
The email-marketing leader changed even though both brands were widely visible.
For a growing 20-person business, the API recommended Brevo on six of eight questions and ActiveCampaign on five. ChatGPT Search recommended ActiveCampaign on all eight and Brevo on seven.
API with webAPI diagnosis: use Brevo as the primary competitive benchmark.
ChatGPT SearchSearch diagnosis: ActiveCampaign holds the strongest sampled position.
“Both dashboards can contain correctly counted data while creating different competitor priorities.”
The trust leader changed for a high-consideration mortgage purchase.
The web-enabled API recommended Better Mortgage on three privacy-and-trust questions and Rocket Mortgage on one. ChatGPT Search recommended Rocket on four and Better on three.
API with webAPI diagnosis: Better owns the strongest association with privacy and trust.
ChatGPT SearchSearch diagnosis: Rocket owns the strongest sampled association.
“That difference can redirect messaging research, digital PR, comparison pages and sales battle cards.”
These examples do not prove which company a consumer would ultimately choose. They show that the two collection methods would not produce the same market diagnosis or action list.
Finding three
Local businesses had the largest measurement problem.
Local answers can be sensitive to the city, precise location, review sources, directories, opening hours and whichever pages are retrieved at that moment. The two methods also received location context differently. This pilot cannot identify one cause, but it shows why a method validated on software brands should not automatically be assumed reliable for dentists, plumbers or moving companies.
Web search improved the API—but did not make it interchangeable.
Without web search, the API disagreed with ChatGPT Search about sampled-business visibility in 21.8% of comparisons. With web search, that fell to 16.4%. When both methods produced one unambiguous panel winner, the exploratory different-winner estimate fell from 41.7% to 22.9%.
This is not an argument against APIs. A scalable API can be appropriate for frequent directional monitoring. A smaller customer-facing panel can then calibrate whether that scalable method still represents the experience it is intended to measure.