dimanche 1 février 2009

4.3.5 Others search engine Internet coverage

“This part needs additional information and improvements and is then not finished
yet.”
You have a last category of search engine which are the one coming from the
semantic web concept which are not based on keywords requests. I found one of
those which is called « Who is like it » and gives as results websites similar to the
one you just gave him.


Illustration 13: “Who is like it” search engine

Of course solutions have been found in order to create the perfect search
engine including all those technologies. But as you can imagine it is not an easy
thing if firstly search engines did not take the same standards as Boolean operators.
Secondly the more search engine you combine the more results you will get and it
then very complicated to decide which one is more pertinent than the other.
A good example of it is the meta search engine called Dogpile which is quite
popular in the United States it combines the results of 4 big search engines which are
Google, Yahoo, MSN live search and Ask. The main problem is already showed on
the advanced search web page of Dogpile, few are the Boolean operators limited to
4.

Illustration 14: Dogpile meta search engine

The second issue which is relevant is how to decide which results from those
search engines are the most relevant. I showed you previously that when making a
comparison between Google and Yahoo I found what I wanted on Yahoo on the 8th
position. So it is hard to decide which result is the most relevant and to this game you
are the only judge of the situation.
So the hypothetical search engine should have the following characteristics:
including all the technologies ever created regarding search engines (Google's
knowledge+Yahoo knowledge+MSN knowledge+...) in order to define the most
powerful algorithm ever. All those companies have of course different objectives
than being the most rational search engine all are looking to be the number one. This
day will may come but it will not be for sure for tomorrow so waiting for it we
should understand how to use one by one all those technologies.
The solution being then to test the search engines one by one but
before testing all of them you have to know that they exist and how to use them but you should at least to know that they exist.

4.3.6 A concrete representation of the World Wide Web

In order to make this work easier a Japanese company set up regularly a web
map of the most famous websites in the world by referencing them by categories this
map could of course be improved but should be a good start:Illustration 15: A representation of the most visited websites

This map is available at the following address: http://informationarchitects.jp/
start/ with all the links included towards the websites it is composed of. It is very
interesting in order to break the search engine dependency phenomenon. On this map
are located all the most famous websites for 2008 you can then see all the most
influential websites and where to seek for information. For example if I want to look
for a video I may follow the gray line and test all the websites which are on it to find
the video I am looking for.
In order to conclude this part which actually should not be the core of the all
thesis the Web is big and in order to discover it you need to know how to.
We could compare it to the real world when living in country A you receive as
feedback from country B by different ways (people who are moving from country B
to A, the news, the books and documentations you have about country B) but you
will never be physically in country B and for this there are some information that you
could never get. Of course you can get nearer to those information by crawling more
and more the web with your country A search engine (it will be like documenting
yourself more and more about country B) but it will never be like being and living in
country B.

4.4 The gap between search engine dependency and data quality

“This part needs additional information and improvements and is then not finished
yet.”
In this part I will explain how to put in evidence the gap of information
between being search engine dependent of only one search engine and using the most
rational tool to search for information on the Internet.
It may not be easy to prove it concretely so I may need to prove it by making
empirical studies.
I succeed to put in evidence so far that:
• Internet users are search engines addicted;
• Internet users have few knowledge about search engines awareness;
• I have now to prove that this is bad;
My first point will be to show that search engines are using different
technologies provide different results.
Here is a comparison I made for three search engines through
http://www.thumbshots.org/Products/Thumbshots/Ranking.aspx which shows how
many similarities search engines have among them. Here I said if I type in
“Universität Kassel” what are the results that those search engines have in common
(in blue). And I moreover added the option to highlight me the website of the
university of Kassel (in red) which for me is a sign of relevancy of my request.

Illustration 16: Comparison among Google, Yahoo and MSN
Here as we can see:
• Google has 7 similarities with Yahoo out of 60;
• Google has 10 similarities with MSN out of 60;
• Google found two times the website I was looking for;
• Yahoo has 4 similarities with MSN out of 60;
• Yahoo found one time the website I was looking for;
• MSN found 4 times the website I was looking for;
This analysis is then confirming what I was supposing before search engines do not
look for information at the same place.
Then one could ask about the pertinence of the results, is it worthwhile to display
more than one time on one page the website of the university of Kassel? Or is it a
sign of relevancy? So here we have an interpretation according to the Internet user.

One thing is sure search engines using different technology provide
different results.

Chapter 5: The Google example

“This part needs additional information and improvements and is then not finished
yet.”

Google is for sure in some parts of the world the best example we can find of
the search engine dependency phenomenon I described.

5.1 Google

In January 1996 a 24 year-old PhD student called Larry Page studying at the
University of Stanford was looking for a theme for his thesis. Encouraged by his
supervisor he studied the following topic “exploring the mathematical properties of
the World Wide Web“ working in collaboration with another student called Sergey
Brin. To make it simple, it is from this work and collaboration which will came up
“Google Inc” (officially created in September the 7th 1998).
Two months later Google is already included in the Top 100 of world
websites of PC magazine (a reference in the United States for computers).
Even if Google is formerly a web based application in English it is a worldwide
service available on the Internet for all. As his creator (Larry Page) said "Google's
search engine has always had strong global appeal"35.
Google is nowadays the most famous search engine. In ten years as the
Millward Brown report said36 Google will become the most powerful brand in the
world.

5.2 Google's success

The broadest explanation I found is the following « Google provides for free
a useful service that people actively seek out »37. And when you think about it it is
definitely true. People are going on Google because they are all looking for that kind
of services. But how can it be more successful than its fellows? Here I could say that
in general Google is giving better results which may be one reason, but moreover it
is providing added services such emails, blogs, news, Advertisement programs.
Moreover Google's health as a company has been well preserved by making
good choices when making acquisitions and mergers, they internationally developed
themselves very well.

5.3 Google dependency state

If we take Europe we can see that it is definitely a Google dependent
continent:
Illustration 17: Google's domination in Europe
As we can see here Google (in blue) is not only the most used search engine
in those countries it is in fact like the only search engine present at the continental
level.