“This part needs additional information and improvements and is then not finished
yet.”
Let's imagine that you and I just performed a request under a specific search
engine and that the answers given do not satisfy you entirely. You found the
information but it is not developed enough, not signed, too old...
dimanche 1 février 2009
4.2 What the search engines do not tell you
One of my former English teacher taught me one day that there is different
ways of not giving information, one is to lie and one is to not say the information. I
don't think search engines are lying and cheating even if in the case of Baidu the
Chinese search engine there are high suspicions on it28. Some are saying that in order
to be well ranked you need to pay Baidu for it and it has been said as well that the
Chinese government as a big role to play when displaying the results.
The recent Milk scandal in China (Baidu accepted to high ranked unlicensed
companies which were providing fake milk in exchange of money) showed how
dangerous can be a poor data quality search29.
I however have the proof that search engines are not telling you everything
when displaying results and that this rule is general for all the major search engines.
In Google case, search engines are censored according to the country in
which you are making your research from30 and it touches all the country around the
world even countries such as France and Germany.
Most of the time those censorship acts are for your welfare or to protect
national security interest.
ways of not giving information, one is to lie and one is to not say the information. I
don't think search engines are lying and cheating even if in the case of Baidu the
Chinese search engine there are high suspicions on it28. Some are saying that in order
to be well ranked you need to pay Baidu for it and it has been said as well that the
Chinese government as a big role to play when displaying the results.
The recent Milk scandal in China (Baidu accepted to high ranked unlicensed
companies which were providing fake milk in exchange of money) showed how
dangerous can be a poor data quality search29.
I however have the proof that search engines are not telling you everything
when displaying results and that this rule is general for all the major search engines.
In Google case, search engines are censored according to the country in
which you are making your research from30 and it touches all the country around the
world even countries such as France and Germany.
Most of the time those censorship acts are for your welfare or to protect
national security interest.
4.3 The best way to get data quality
Here is maybe the most interesting part of my work which consists in giving
the solution to the main issue I highlighted since the beginning.
Until now I just raised questions regarding the risks of dependency and not
showed what you should do in order to explore the web properly.
the solution to the main issue I highlighted since the beginning.
Until now I just raised questions regarding the risks of dependency and not
showed what you should do in order to explore the web properly.
4.3.1 The sub-search engines
As I said previously Google is too big, the bigger it is the more precise your
request have to be. Few people knowing the existence of Boolean operators it then
comes more and more difficult to get the right information. This is why in fact
Google put at your disposal some sub-search engines in order to make your research
easier, for example: Google books, Google videos, Google images. We all agree that
looking for pictures could be done on the research bar of Google, but it is far more
convenient to make it directly from « Google images » because it is displayed better.
The main problem is what I described previously: « search engine awareness
within a search engine ». Am I aware that Google has the following services?
I could sum up it all in a scheme:

page. He is then aware here of those specialized search engines in order to make
more precise research on the following information: images, maps, news and
products. It is however under his own initiative to discover what is going on under
those search engines and to discover them. I do not have a proof of what I am saying
but I may presume that in order to buy a product more people are going on Google,
type eBay or Amazon and go and buy on those websites rather than going on Google
products. Google products is however including all those websites in his database so
it should more convenient to have a look on Google products first in order to get the
best price rather than going individually on eBay and Amazon.
We can then symbolize the Internet users awareness of Google search engines
according to this scheme.
The more we get into deep of those levels and the least the Internet user is
aware of it.
Level 2 is still available on the first page but with two clicks.
Level 3 is no more available on the home page but can be accessible in three
clicks.
Level 4 is even not accessible from Google main website and that you have
too be aware of it in order to access it.
Level 5 is hypothetical and constitute the unknown Google projects, often
developed under another name.
All those Google sub levels have been created in order to make easier
research on a specific request. When looking for an image it is far
more convenient to pass through Google images rather than the general Google.
request have to be. Few people knowing the existence of Boolean operators it then
comes more and more difficult to get the right information. This is why in fact
Google put at your disposal some sub-search engines in order to make your research
easier, for example: Google books, Google videos, Google images. We all agree that
looking for pictures could be done on the research bar of Google, but it is far more
convenient to make it directly from « Google images » because it is displayed better.
The main problem is what I described previously: « search engine awareness
within a search engine ». Am I aware that Google has the following services?
I could sum up it all in a scheme:

Illustration 9: Some sub-search engines of Google
For a Google search engine dependent everything starts from Google homepage. He is then aware here of those specialized search engines in order to make
more precise research on the following information: images, maps, news and
products. It is however under his own initiative to discover what is going on under
those search engines and to discover them. I do not have a proof of what I am saying
but I may presume that in order to buy a product more people are going on Google,
type eBay or Amazon and go and buy on those websites rather than going on Google
products. Google products is however including all those websites in his database so
it should more convenient to have a look on Google products first in order to get the
best price rather than going individually on eBay and Amazon.
We can then symbolize the Internet users awareness of Google search engines
according to this scheme.
The more we get into deep of those levels and the least the Internet user is
aware of it.
Level 2 is still available on the first page but with two clicks.
Level 3 is no more available on the home page but can be accessible in three
clicks.
Level 4 is even not accessible from Google main website and that you have
too be aware of it in order to access it.
Level 5 is hypothetical and constitute the unknown Google projects, often
developed under another name.
All those Google sub levels have been created in order to make easier
research on a specific request. When looking for an image it is far
more convenient to pass through Google images rather than the general Google.
4.3.2 The size of the Internet
How big is the Internet? Is Google indexing all websites?
It is really hard to answer to those questions but however possible to set
up some estimations according to some information31. Those sources are
saying that in 2005 the size of the Internet is estimated to 5 million terabytes and
Google's index to 170 terabytes which would mean that Google is processing only
0,000034%. However the Internet is also containing what we call the invisible web
composed of websites that owners do not want its content indexed as well as
websites which are protected by a password. In 2004 this invisible web was
estimated to be 500 times bigger than the visible web. It has been said as well that
Google is indexing invisible web only recently32.
A clever calculation will then give us:
Internet size = Visible Web + Invisible Web;
Internet size = 501 * Visible Web;
Visible Web = Internet size / 501;
Visible Web = 5 000 000 / 501;
Visible Web = 9980 terabytes;
Google Index = 170 / Visible Web;
Google Index = 1,7%
This estimation is of course only an estimation and could be full of errors. I
however find it more useful than no information at all. I would also emphasize the
fact that Google has no interest in indexing bad quality websites and that
technologies have evolved in the last few years and that this rate should be of course
far higher than those 1,7%. Whatever is the final result my point is the following:
Google is not the Internet and is not processing all the web... but does
all the web need to be indexed?
It is really hard to answer to those questions but however possible to set
up some estimations according to some information31. Those sources are
saying that in 2005 the size of the Internet is estimated to 5 million terabytes and
Google's index to 170 terabytes which would mean that Google is processing only
0,000034%. However the Internet is also containing what we call the invisible web
composed of websites that owners do not want its content indexed as well as
websites which are protected by a password. In 2004 this invisible web was
estimated to be 500 times bigger than the visible web. It has been said as well that
Google is indexing invisible web only recently32.
A clever calculation will then give us:
Internet size = Visible Web + Invisible Web;
Internet size = 501 * Visible Web;
Visible Web = Internet size / 501;
Visible Web = 5 000 000 / 501;
Visible Web = 9980 terabytes;
Google Index = 170 / Visible Web;
Google Index = 1,7%
This estimation is of course only an estimation and could be full of errors. I
however find it more useful than no information at all. I would also emphasize the
fact that Google has no interest in indexing bad quality websites and that
technologies have evolved in the last few years and that this rate should be of course
far higher than those 1,7%. Whatever is the final result my point is the following:
Google is not the Internet and is not processing all the web... but does
all the web need to be indexed?
4.3.3 Single search engine Internet coverage
“This part needs additional information and improvements and is then not finished
yet.”
Here is a more optimistic representation of Google Internet coverage I made
in order to show that Google is not the Internet and that even within Google's sphere
Internet users cannot access to all the information:

Google dependent people are not only using Google when making research but as
well Google partners all symbolized by the sign:
Powered by Google means according to an IT company called Alacra33:
Alacra uses Google Search Appliances to create the Alacra Compliance Web. The
use of Google Search Appliances combines the power of Google search technology,
including the ability to find the highest quality and most relevant documents, with
Alacra's domain expertise in selecting those web sites and pages which are relevant
for AML compliance.
Which means that here you have a partnership working more or less in the same way
that using a Google specialized search engine.
There are thousands and thousands search engines all specialized in a specific
field on the Internet which are using Google technology to search sites such as
Tourism, tutorials...
By using all those specialized search engines you are crawling better the
Google's coverage space.
As you can see here on this configuration by using specialized search engines
powered by Google you will always browse Google's cyberspace and when we know
that Google is not the Internet we can then ask ourselves how we can get the best out
of Internet.
On January the 11th I went on both websites http://www.aol.fr/ and
http://www.google.fr/ and type the following request 'les moteurs de
recherches' both results on the first page were identical. The only
difference is that Google gave me 3,210,000 results and AOL 290,000
(9% of Google results) so as I developed above AOL is looking at the same place as
Google but is applying more filters. Is it really useful considering that AOL is a
general search engine? The answer maybe be given in their home page at
http://search.aol.com/aol/webhome « The AOL Search engine delivers great search
results, enhanced by Google, plus relevant multimedia results delivered on a single
page-so you can search less and discover more . »34
If you go on http://www.google.com/coop/cse/ you will have a good example
of the search engine powered by Google. You can even create your own one. All is
done by using Google technology and you are just applying your own filter by listing
the websites where you want Google to look inside.
yet.”
Here is a more optimistic representation of Google Internet coverage I made
in order to show that Google is not the Internet and that even within Google's sphere
Internet users cannot access to all the information:

Illustration 10: Google's Index coverage
Google dependent people are not only using Google when making research but as
well Google partners all symbolized by the sign:
Powered by Google means according to an IT company called Alacra33:Alacra uses Google Search Appliances to create the Alacra Compliance Web. The
use of Google Search Appliances combines the power of Google search technology,
including the ability to find the highest quality and most relevant documents, with
Alacra's domain expertise in selecting those web sites and pages which are relevant
for AML compliance.
Which means that here you have a partnership working more or less in the same way
that using a Google specialized search engine.
There are thousands and thousands search engines all specialized in a specific
field on the Internet which are using Google technology to search sites such as
Tourism, tutorials...
By using all those specialized search engines you are crawling better the
Google's coverage space.
As you can see here on this configuration by using specialized search engines
powered by Google you will always browse Google's cyberspace and when we know
that Google is not the Internet we can then ask ourselves how we can get the best out
of Internet.
On January the 11th I went on both websites http://www.aol.fr/ and
http://www.google.fr/ and type the following request 'les moteurs de
recherches' both results on the first page were identical. The only
difference is that Google gave me 3,210,000 results and AOL 290,000
(9% of Google results) so as I developed above AOL is looking at the same place as
Google but is applying more filters. Is it really useful considering that AOL is a
general search engine? The answer maybe be given in their home page at
http://search.aol.com/aol/webhome « The AOL Search engine delivers great search
results, enhanced by Google, plus relevant multimedia results delivered on a single
page-so you can search less and discover more . »34
If you go on http://www.google.com/coop/cse/ you will have a good example
of the search engine powered by Google. You can even create your own one. All is
done by using Google technology and you are just applying your own filter by listing
the websites where you want Google to look inside.
4.3.4 Multiple search engine Internet coverage
“This part needs additional information and improvements and is then not finished
yet.”
Let's go now deeper in our analysis by taking in account search engine which
are not using Google technologies. All developed their own technology and are
sometimes better than Google in some fields worst in the other. We will then have
something like that:
So here is my point the most the search engines are different and the most it
gives you the possibility to discover the web. Here I just put the main actors but we
have also to consider that it exists a lot of small search engines which developed their
own index and have then their own way to process data.
Of course one could say that there is no interest for an European person to
process information on a Chinese or even Russian search engine because Chinese
search engine should of course be better in looking for Chinese information rather
than an European one. It is definitely true if we are speaking about contents such as
texts, but when it is dealing with pictures or video the reality should be totally
different.
On the other hand the day where this European person is looking for Chinese
information he should then be aware that he should use the Chinese search engine
rather than the Chinese version of the European one. And as we can imagine all the
other players: Yahoo, Baidu, Mail.ru have their own Boolean operators and their own
sub search engines. So it comes back to the idea of search engine awareness. The
more you know about them and the most you are increasing your knowledge about
how to get the best out of Internet.
yet.”
Let's go now deeper in our analysis by taking in account search engine which
are not using Google technologies. All developed their own technology and are
sometimes better than Google in some fields worst in the other. We will then have
something like that:
So here is my point the most the search engines are different and the most it
gives you the possibility to discover the web. Here I just put the main actors but we
have also to consider that it exists a lot of small search engines which developed their
own index and have then their own way to process data.
Of course one could say that there is no interest for an European person to
process information on a Chinese or even Russian search engine because Chinese
search engine should of course be better in looking for Chinese information rather
than an European one. It is definitely true if we are speaking about contents such as
texts, but when it is dealing with pictures or video the reality should be totally
different.
On the other hand the day where this European person is looking for Chinese
information he should then be aware that he should use the Chinese search engine
rather than the Chinese version of the European one. And as we can imagine all the
other players: Yahoo, Baidu, Mail.ru have their own Boolean operators and their own
sub search engines. So it comes back to the idea of search engine awareness. The
more you know about them and the most you are increasing your knowledge about
how to get the best out of Internet.
Inscription à :
Articles (Atom)

