> If it were just click data, how would they get the terms?
Exactly. It's not "click data" at all. It is monitoring user behavior on search engines, using both the clicks and the queries. Maybe it's not monitoring just Google as a search engine (although we have no proof of that yet: it seems it's just watching Google) but given Google's market share in search it doesn't make much difference.
From the article:
> I don’t even work in search and I could spot the real situation
This sentence is at the same time arrogant and funny. He doesn't work in search, but he's certain he's spotted "the real situation". How did he fact check it, besides asking himself if he was correct and answering "yes, obviously I'm right. I'm always right -- and I don't even know anything! I amaze myself."
> certain he's spotted "the real situation". How did he fact check it
By reading the official MS blog post[1], watching the Farsight conference live, and being confident in my own judgements based on the observed evidence. No other conclusion is supported by the evidence - you say "It is monitoring user behavior on search engines" - where's the evidence that it's exclusively search engines they're monitoring? Let alone exclusively Google? There has been no evidence yet presented. Therefore all we can conclude from Google's sting is that Bing use URL/click data. Which is exactly what I said in the original thread[2], in my blog post, and lo and behold it's what MS later said. That's how I know it's the real situation.
[1] like where Harry says "A small piece of that is clickstream data we get from some of our customers, who opt-in ... To be clear, we learn from all of our customers".
You checked it against the assertion of the most interested party?
It's not "click" data, it's the correlation between search term and SERP. That someone visited such and such a page after a click isn't the issue. That someone Googled for a unique term before clicking is at issue.
> You checked it against the assertion of the most interested party?
There are currently exactly two sources: The Google post and the MS post. Who are more likely to know about what MS are doing? MS. I check my theory about what MS are doing with the source most likely to contain correct information.
> It's not "click" data
It's 'clickstream' data (MS's term). A 'click' comes from a page and goes to a page. That's the data MS were capturing. The page the click happened on (query happens to be included in URL), and the page the click went to. It's click data.
Your assertion Bing is most likely to be accurate about Bing ignores self-interest and spin.
However, agreed -- clickstream means series of clicks, and the actual data is a series of URLs.
The query "happening" to be in the URL has no "search > result" meaning without a parser being told to look for Google's particular keyword query indicators and correlate the subsequent page. As most URLs are not searches, this is not emergent behavior; it's programmed.
People also talk about this being a "weak" signal, but given search volume (or clickstream volume if you prefer) on Google versus other sources, even if this code is generic (e.g., recognize all "q=blah" or "search=blah" as keywords and correlate the following URL), it seems the signal would be strong indeed. Google's weak signal would provide several times more correlative data to Bing than Bing's own clicks.
Not that there's anything wrong with that! But Bing's blog assertions feel disingenuous -- they play this game well:
> Your assertion Bing is most likely to be accurate about Bing ignores self-interest and spin.
We can't apply skepticism to one source and not the other. Either Google and Bing are not blogging with self-interest and spin - and thus Bing are more trustworthy because they're blogging about themselves, or they both are blogging with self-interest and spin - and still Bing are more trustworthy because they're blogging about themselves.
You just can't legitimately discount what Bing say because of self-interest and spin without also discounting what Google say for the same reasons.
> The query "happening" to be in the URL has no "search > result" meaning without a parser being told to look for Google's particular keyword query indicators and correlate the subsequent page.
No. remove non-alpha from entire URL with no preconception about search queries or any of that. You're left with "google com search q QUERYTERM". All the words apart from QUERYTERM has plenty of other signals in Bing's system. If QUERYTERM is a highly unusual word then all Bing have to go on is the data they gleaned from Google.
> We can't apply skepticism to one source and not the other. Either Google and Bing are not blogging with self-interest and spin - and thus Bing are more trustworthy because they're blogging about themselves, or they both are blogging with self-interest and spin - and still Bing are more trustworthy because they're blogging about themselves.
Libel laws prohibit indiscriminate accusation, while there's no law against puffery.
The premise a company's own public relations messaging is more trustworthy than an outsider because the company's PR is about themselves seems without merit -- otherwise we would deem companies more trustworthy whenever an outsider points fingers, and send all the journalists, whistleblowers, and wiki-leakers home. "Nope, sorry, I believe the company, because they're talking about themselves."
Google presented incontrovertible data. Bing's PR tactic is "Google does this or that worse thing and profits off it" -- distracting hand waving -- "plus we're not copying anyway" -- deliberately disingenuous.
Responsive would be:
Of course our toolbar is recognizing search terms across the top N search sites, and correlating human selected results as an indicator of search intent and result quality for those search terms. This is the same thing you do when you look at your own web stats and check inbound search terms for your own pages: 'How relevant are my pages, and am I showing my users what they are looking for?'
This is the very definition of 'improving your search experience' as outlined when you install our toolbar. We're thrilled so many of you chose our Internet Explorer browser and Bing Toolbar that this provides us meaningful data on user search intent. We want to thank Google for demonstrating we are truly 'improving your search experience' using well accepted Internet crowd-sourcing techniques.
We agree, however, that generating correlations solely from competitor listings -- when we have no existing correlation in our own data corpus -- could be misperceived, so going forward, we will not create correlations solely from competitor results where none existed in our data. However, like every webmaster, we will continue to use crowd-sourced search term and results data from across the web to refine our suggestion order towards best predicting the information you want to find.
> The premise a company's own public relations messaging is more trustworthy than an outsider because the company's PR is about themselves seems without merit
> where's the evidence that it's exclusively search engines they're monitoring? Let alone exclusively Google? There has been no evidence yet presented
The evidence that has been presented by Google does show that MS is monitoring Google's search results, or at the very least users' behavior when using Google.
I agree that there's no evidence that Google is the only search engine that's being monitored, but it makes little difference since for all practical purposes Google == search.
There's also no evidence that MS monitors other websites / behavior other than just search engines, but it's unclear how that would work? You can't just associate two websites because some users go from one to the other (correlation vs. causation, etc.)
- - -
At this point we're still in the "he said / she said" phase, but Google has more evidence and MS is coming out as incredibly defensive (e.g., raising an inquiry from the European Commission[1]: what does that have to do with anything!!?!)
> The evidence that has been presented by Google does show that MS is monitoring Google's search results, or at the very least users' behavior when using Google.
I said "exclusively".
> there's no evidence that Google is the only search engine that's being monitored
There's no evidence that it's only search engines.
> You can't just associate two websites because some users go from one to the other
sure you can, if they go via a link. Now you can just notice the link and say "great, there's an association", or you can monitor which links people click on and weight the graph accordingly. "great, there's an association, and this link is more popular than that one". Nice data to capture. Not constrained to search.
Google (and pretty much every other search engine) puts your search terms right up in the page title, and they show up in several other places around the page. The Bing scraper couldn't miss them.
Honestly, if Bing isn't giving Google special treatment, this whole debacle shows that their click tracking model works.
Exactly. It's not "click data" at all. It is monitoring user behavior on search engines, using both the clicks and the queries. Maybe it's not monitoring just Google as a search engine (although we have no proof of that yet: it seems it's just watching Google) but given Google's market share in search it doesn't make much difference.
From the article:
> I don’t even work in search and I could spot the real situation
This sentence is at the same time arrogant and funny. He doesn't work in search, but he's certain he's spotted "the real situation". How did he fact check it, besides asking himself if he was correct and answering "yes, obviously I'm right. I'm always right -- and I don't even know anything! I amaze myself."