Affichage des articles dont le libellé est Master thesis: Intermediate report. Afficher tous les articles
Affichage des articles dont le libellé est Master thesis: Intermediate report. Afficher tous les articles

dimanche 1 février 2009

Rapport intermédiaire de thèse

Salut à tous,

Comme promis je vous publie mon rapport intermédiaire de thèse.
Pour le télécharger cliquez sur le lien suivant (mais il vous faudra un compte gratuit sur slideshares):
Lien pour la thèse
Le rapport final est prévu pour juin.
Bonne lecture.

Risks of search engine dependency and its influence on data quality

Thesis intermediate report submitted for the European Master in Business Studies
(EMBS)
by Ronan CHARDONNEAU
Institut de Management de l'Université de Savoie d'Annecy (FR)
Università degli studi di Trento (IT)
Universität Kassel (GER)
Universidad de León (SP)
Date of submission: 26th January, 2009
Master Thesis

Contents

Foreword.......................................................................................................................6
Chapter 1: Introduction of the topic background..........................................................8
1.1 Relevance of the subject...................................................................................10
1.2 Major terms......................................................................................................11
1.3 Focus, goals and structure of the report...........................................................11
Chapter 2: Concept of data quality.............................................................................13
2.1 Data quality definition......................................................................................14
2.2 The importance of data quality.........................................................................15
Chapter 3: Search engines dependency.......................................................................16
3.1 Search engine market configuration.................................................................17
3.1.1 Search engine categories..........................................................................17
3.1.2 Search engine market...............................................................................19
3.1.3 The search engines in the world...............................................................19
3.1.4 The search engine market shares per country...........................................22
3.1.5 The search engines competition...............................................................23
3.1.6 The semantic web.....................................................................................24
3.2 Search engines dependency aspect...................................................................25
3.2.1 Search engines dependency proves..........................................................25
3.2.2 Search engines dependency aspect...........................................................27
3.3 Search engines dependency problems..............................................................28
3.3.1 Privacy issues...........................................................................................29
3.3.2 Looking for other search engines.............................................................30
3.3.3 Search engine awareness..........................................................................30
3.3.4 Other search engines existence awareness...............................................32
3.3.5 Less confident regarding other search engines.........................................33
3.3.5 Less confident regarding other search engines.........................................33
3.3.6 Even the best cannot provide you everything...........................................34
Chapter 4: Risks of search engines dependency and its influence on data quality.....35
4.1 The information has been found but is poor....................................................36
4.2 What the search engines do not tell you...........................................................36
4.3 The best way to get data quality.......................................................................37
4.3.1 The sub-search engines.............................................................................37
4.3.2 The size of the Internet.............................................................................38
4.3.3 Single search engine Internet coverage....................................................39
4.3.4 Multiple search engine Internet coverage.................................................42
4.3.5 Others search engine Internet coverage....................................................44
4.3.6 A concrete representation of the World Wide Web...................................46
4.4 The gap between search engine dependency and data quality.........................47
Chapter 5: The Google example.................................................................................50
5.1 Google..............................................................................................................51
5.2 Google's success...............................................................................................51
5.3 Google dependency state..................................................................................52
5.4 Google functions..............................................................................................52
5.5 Google added functionalities............................................................................53
5.6 Google success is his weakness.......................................................................53
5.7 Google's disappearance hypothesis..................................................................54
Conclusion..................................................................................................................55
Declaration..................................................................................................................56
List of literature...........................................................................................................57
Afterword....................................................................................................................61

Foreword

As most of the students who has a computer one of my first move when I
wake up is to switch on the computer and to spend my first twenty minutes of the day
on the Internet.
From there I have a look at the last news, I check my e-mails and eventually
exchange some few words with a couple of friends by using online chat applications.
I also check my other email account as well as my blogs and analyze the traffic I got
during the last few days, to finish this process I consult my advertisement account to
see if I got some revenues. I often use as well search engine to look for information
which just came up into my mind during the night.
In the paragraph you just read was the description of my morning routine on
Internet. There is nothing special except that most of the moves I described above are
in fact done on two to three major search engines: Google, Yahoo and Microsoft.
I hardly ever use Yahoo or Microsoft for search purpose but Google is for
sure the website I visit the most to crawl the web but... is Google the Internet?
I got the idea to write about: « Risks of search engine dependency and its
influence on data quality » not because I was using all those Google applications
everyday and was scared about what will happen if I get in troubles with Google
such as privacy issues or if Google just closed. I just write about it because one day I
found Google results not accurate enough.
And from this observation a lot of questions came to my mind:
· Is it me who is not good enough at performing research on the Internet?
· Is it because no one wrote about the information I am looking for?
· Is it because the information is not on the first pages in Google that I have to
browse all the pages in order to find it?
· Is it because Google is not good enough?
· Is it because the information is hidden in some other documents such as PDF,
pictures, videos?
· Is it because I have to use another way to crawl the web and if yes how?
You see here how a simple observation can raise a lot of questions.
I hesitated a lot about writing on this topic, the main problem I got was that I
was not convinced that there is a potential risk of being search engine dependent. The
reason is that companies such as Google are working hard in order to fit Internet
users expectations and the vision we get is that they are doing a wonderful work. The
problem is that there could be a difference between perception and real facts and this
is exactly what I am eager to discover here.
Can we measure how huge is the gap between the information we were
looking for and the one of search engines as Google are providing us?
Search engines are set up to find information on the Internet, information
being the basis of any good decisions making we can then understand how important
and interesting it is to write on this topic.
I hope you will appreciate this reading as much as I did when making my
research.

Chapter 1: Introduction of the topic background

I will not surprise you if I say that Internet has been created to share
information and to communicate with each others.
It is hard to evaluate how big is the Internet, estimations among companies
are very different, it varies from 15 to some 30 billion Web pages1. The number of
websites is increasing everyday and estimated at 185,167,8972 with a constant
augmentation since the creation of the world wide web.


Illustration 1: Total Sites Across All Domains August 1995 - January 2009

Habits have changed since the creation of the Internet and websites are used now in
diverse manners if it comes to be a standard for companies (recognized as a mark of
trust, seriousness and quality) it is also a space for many individuals (blog
phenomenon). As an example regarding France, in June 2008 14% of French people
above 12 year-old which means 22% of French Internet users are authors of a blog or
a website3.
The banalization of the Internet and the fact that anyone can create his own
website for free increase the feeling we have regarding the Internet: a true jungle of
information and even sometimes real “dump” regarding information accuracy.
Websites can be accessible through three channels:
· Direct access (for example you know the website address by heart, you put it
in your favorites or you find a website on a business card and you are typing
it in the address bar);
· External links (you access to a website which has the link of another
website, this is the case in most of websites, catalogs, advertisement);
· Through Search Engines (you use a dedicated application by typing in some
keywords in order to get suggestions of what you are looking for);
As you can see from this list if you use only the first two ways to crawl the
web it comes to be too rigid and not wide enough. It has been said as well that the
first way is disappearing more and more in profit of search engines4.
So one could say that there is currently two main ways to crawl the web, from
link to link and by using search engine.
This last one being indispensable in order to crawl the web properly.
More and more information are put on the Internet which makes it
come a true jungle. The only way to crawl those information properly
is to use search engines.

1.1 Relevance of the subject

Internet is becoming more and more our information provider. "In 2002 a
study from the U.S. Department of Commerce estimated that 36% of the American
public over the age of 3 used the Internet to search for product and service
information in 2001. This usage represented a substantial increase in information
search behavior from 2000, when 26% of the American public reported using the
Internet to search for product and service information."5
Between 2000 and 2008 the USA got an increase of Internet users of 130,9
%6 we can then imagine how important is searching information through Internet
nowadays.
The number of Internet users is estimated to 1,463,632,361 (world population
6,676,120,288) with a growth rate from 2000-2008 fixed at 305.5 %7.

1.2 Major terms

In this thesis you will hear a lot about the following terms: search engines,
search engine dependency and data quality.
Search engine is the most flexible technology which has been created in
order to crawl the web. A search engine is no more than a web application which is
processing data. A search engine does not create data it just process some
information it has in his index.
I describe search engine dependency as the fact that one does not make the
choice between one search engine and another. Search engine dependency is very
relevant in most of the countries. Most of the people are using only one single search
engine when performing requests on the web. I will speak as well about bad habits
when dealing with search engines. For example you can be dependent of using a
search engine but using it badly.
Data quality is the quality of data. Data are of high quality "if they are fit for
their intended uses in operations, decision making and planning ". Alternatively, the
data are deemed of high quality if they correctly represent the real-world construct to
which they refer. These two views can often be in disagreement, even about the same
set of data used for the same purpose8.

1.3 Focus, goals and structure of the report

The focus of this work is to show if there are some risks of being search
engine dependent and if there are some, is the gap of information significant
between a search engine to another?
Goals are numerous:
· to put in evidence how bad it is to be dependent of an unique search engine;
· to show how important it is to have reliable information;
· to show how big is the gap between a general search engine and a specialized
one;
· to look for alternatives if tomorrow the main search engines disappear;
· to show the weaknesses of general search engines;
· to show and discover other search engines and put in evidence their
usefulness;
· to show that general search engines are not well used;
· to discover the future expectations people have regarding search engines and
what are the coming technologies in this field;
The structure of the report should be done as follow:
The first idea is to introduce the concept of data quality. We all go on
search engines in order to find information whatever it is or at least to find the
answer to one of our question (how much does a Paris-Berlin train ticket costs? What
is the weather in New York? What is the last result of our favorite soccer team?).
The second point is about Data quality which is very interesting in order to
understand why search engines exist.
The next point is dealing with the world of search engines and the
dependency which is outcome from them.
Analyzing the world of search engines is fundamental to understand how
Internet is not as rational as we could think (search engines may be not the Internet,
search engines may be different from a country to another).
Then we will have a look on figures which show quite clearly that people are
not using several search engines when making research on the Internet but an unique
one that I will define as the dependency concept.
Once this definition given I will focus on the heart of this work which are the
risks that this dependency provide and its influence on data quality.
Google being in Europe the most used search engine I will use it as a major
example in my last part.

Chapter 2: Concept of data quality

When I firstly decided to write about the concept of data quality for search
engine I did not mean at the beginning that data has to be perfect. I just meant that
when I write a request on any search engine I expect to have as results some
answers which fit to my expectations. The last example I have in mind is when on
the 31th of December a friend of mine looked for the « average height of Korean
people » (request were made on Google) and all the results on the page were dealing
about the « average size of the sexual organs of Korean people ». I have high doubts
that no information regarding the average height of Korean people does not exist on
the web it however has been what Google seemed to tell me.
Here as you can see in this example it raises a lot of issues regarding data
quality. My first expectation regarding data quality was then to get some results
which fit to my request. But looking it deeper what is the value of a supposed correct
answer that I could have found written by someone like you and I (not specialized in
this topic).
In fact here everything depends on how professional the data as been written,
this is what I am explaining in this chapter.

2.1 Data quality definition

“This part needs additional information and improvements and is then not
finished yet.”
Data quality is defined by four criteria:
· Accuracy: it means that the information has to be true so based on real facts.
Here we see the importance of having the source of the document.
· Timely: the data given has to be dated. The most striking example could be
the one with stock exchange, what is the value of a data regarding a currency
change without the date?
· Meaningful: here it comes back to what I was explaining in the introduction.
Does the result fits with my request? What is the color of a frog? The answer
is green, is it useful? Yes but is it complete?;
· Complete: an information can be for sure useful but will be far more useful if
it is complete.
Here is then the minimum vital of data quality.

2.2 The importance of data quality

“This part needs additional information and improvements and is then not finished
yet.”
As I mentioned it there are two levels of data quality:
· Poor quality one: for example you and I are looking for information
whatever it is in order to get a quick answer or even a mere comment. Here
two theories are facing each other:
▪ An information should be given even if it is possibly wrong. It of
course can be useless if the information is a hoax but will it be in the
case of a warning of a terrorist attack?
▪ An information should be given only if it is 100% accurate. Here it is
interesting in order to not create polemical situations for nothing.
There are no rational choice to make here, sometimes you need to use the first
theory and sometimes the other one.
I would describe poor quality data as mass consumption information which
mean good enough for “lambda” people but not that useful for businesses and even
less for researchers. With the increasing of Internet users more and more mass
consumption data are created and of course their number is bypassing the one of
quality ones.
· High quality one: here we find the data which fits the four criteria I
described above. High quality data are useful in order to take complex and
rational decisions.

Chapter 3: Search engines dependency

In order to understand well this part we should have firstly a look at the
search engine market configuration. Even if for a lot of people search engine is a
synonym of Google it may be not true for some other people. I made a strong effort
during this work to make it as global as possible. Europe is good as well as the
United States but we will never got the all answers of our issues if we always
consider these two parts of the world.

3.1.1 Search engine categories

3.1 Search engine market configuration
3.1.1 Search engine categories
The world of search engines is not as uniformed as we could have think. I
until now identified three kind of search engines:
· Standard: the most well known search engines such as www.google.com,
www.livesearch.com (Microsoft) , Ask (http://www.ask.com/?o=312) they are
looking for any kind of information through the Internet and are characterized
by a very light interface:


Illustration 2: A light interface: the homepage of the search engine www.ask.fr

· Portals: may be the most complex search engines to analyze on a figures
point of view. Portals are characterized by a lot of information on their home
page including the search engine function. It is then difficult to make the
difference here between the success of the portal and the success of the search
engine on this web page. The best example I found is the one of mail.ru
which is the most visited website in Russia far before Google. It seems when
looking at the search engine statistics that the research made on Mail.ru are
not accounted or that no one is using the search function of Mail.ru. So it is
sometimes hard to measure the success of some search engines. The most
well known portal is Yahoo.
Illustration 3: Home page of the Mail.ru portal

Specialized search engines: they to my point of view belong to a
subcategory of the first group. The major problem of standard search engines
is that they are too big. Specialized search engines are then no more than a
special function which is working as a filter. It is then far more easier to find
the information you are looking for on a specific website rather than using
standard ones. A good example of it can be the one of Ebay, it is of course far
more convenient to go on the Ebay website and use the search engine directly
from there rather than going on Google and writing a request such as 'Ebay
buying socks'. I mentioned it again but in most of the cases specialized search
engines are not a revolution they are just one part of standard search engine
and are not a new technology by themselves.

3.1.2 Search engine market

The search engine market is segmented by an enormous leader: Google, a
follower who is far from him: Yahoo and a large amounts of small search engines.
The dominant position of Google may stay for years and years and only a technical
revolution could really jeopardized him.


Illustration 4: Top 10: Search websites in the world August 2007
As you can see here I identified what we could call « The big four » which
represents the four major search engines. Then comes three specialized search
engines (7,49% all together) but their success are limited to the size of the website
they are browsing. Then came in last positions what I would called the dead
champions which once have been great and well known search engines but which
are nowadays in decline and will one day probably disappear.

3.1.3 The search engines in the world

The more I study the E-world and the more I realize that the web is exactly
the reflect of our society but in a non-physical aspect. Internet did not break cultural
differences and search engines really show it.
As we just saw Google represents 6 research out of 10 in the world but does it
mean that each country in the world has a population of 60% Google users?
To answer this question let's have a look at the following map:
Illustration 5: The most visited website by country

As we can see here the world is not covered entirely by Google. We clearly
have some Google countries, Yahoo countries, Mail.ru countries and so on and so
fourth.
We can however notice some striking information such as almost all the
American continent is using Google as well as Europe, Northern Africa and
Southern Africa, Australia, India. In one word almost all countries which have strong
links with the Anglo-Saxon culture.
Then comes what I would qualified as a cultural wall which is starting in
Eastern Europe and which is finishing in Russia. Here are the ex-soviet countries. I
unfortunately have no concrete proves of what I am saying but I suppose that there is
a kind of a « boycott of American technologies » and support of Russian
technologies. The recent partnership between Yandex (main search engine in Russia)
and the browser Firefox raised those suspicions10.
Russia is not the only country in this situation, China also. The recent
advertisement broadcast by Baidu (the leader search engine in China) shows clearly
the will that Chinese public institutions are ready to protect their territory.11
As you can see Asia is the region where Google is the least present. It is also
the continent where are gathering a lot of different search engines.
I may not emphasize the diversity of the Caribbean area as well as Center
Africa which are areas where Internet is not that well implemented which mean then
that the battle to take the lead is not finished yet. For example is it really relevant to
say that Yahoo is the leader in Cameroon which is a country with less than 500,000
Internet users?
I however will highlight that the Pacific area which is containing all the
« Tigers » (Taiwan, Thailand, Singapore...) are all in red: Yahoo.
The search engine world is then divided in two parts:
· The Google world: which is composed of all the Anglo-Saxon countries as
well as countries which have strong links with the United States or Great
Britain. It is clearly showed regarding India and Australia. We can at the date
of today still identify two countries which are not under Google control:
Czech Republic and Iceland but actually by looking at the figures and the
forecasts it is just a matter of time12.
· The Asian – Pacific world: Asia is composed of a lot of countries and of
course a lot of cultures. Among them we can identify four players:
◦ Mail.ru which is dominating all the ex-soviet countries;
◦ Baidu which has a total control over China;
◦ Naver, a 100% South Korean product which is the best example that
search engines work by culture;
◦ Yahoo which is leader in all the “Tigers” Asian countries.
Yahoo being an American technology such as Google, how is it possible
that Yahoo is so successful in the Pacific area and not elsewhere? The
reason I found is that Yahoo is a shiny portal and that Asian culture on
Internet recognize a quality website to the number of animations on it13. I will then
add that Japan is a strong pole of Internet with one of the highest rate of Internet
integration in the world per capita14. This is why I think Yahoo is so popular in this
region.

3.1.4 The search engine market shares per country

One of the most complete work I found on this topic after mine is the
« Global Search Report 2007 »15. What stroke me the most in all the countries
studied in this report as well as in all the report and research I made until now are the
search engines market configuration which looks like very often to this:
Illustration 6: Search engine market shares in
Czech Republic
It is very rare to find a country where there is a close competition among
search engines. Even if in the High Technology world things change from a day to
another you have often the following configuration where the first search engine
is leading the game by more than 30 points on its followers.
This is typical from the search engines market or you are adopted by a
population or you are not. This trend seems quite relevant in the
software industry, people seem to look for a standard used by all. This is the case for
the Operating System industry, the browser industry, the e-learning industry. The
explanation I found for the success of search engine within a population is the word
to mouth, this is how Google has been so successful isn 't it? How never heard
sentences such as « you just have to Google it » Google is even nowadays in
dictionaries as a verb16.
As a conclusion I would say that in the world of search engines you are first
or you are nothing.
I have to mention as well that the market has a lot of small local search
engines which are if original enough bought by the big ones or if not will disappear
quickly (some example are coming in the news every month). The only key of the
success on the short term seem to be advertisement but on the long run you need the
technology behind in order to compete.

3.1.5 The search engines competition

Google has been created in 1998 and at that time search engines were already
in place, it did not scared Google and one after the other Google bypassed all of
them. In fact among the big four Yahoo is the oldest (1994) and Microsoft the
youngest (2003). Even if the battle seems to be finished it will take a lot of time to
Google to be the number one in all countries (everything being linked to culture
rather than rationality) which in fact is giving hope to its followers.
At the time I am writing this thesis discussions are still on the way between
Yahoo and Microsoft in order for Microsoft to buy Yahoo search technologies. We
can understand how strategic a such acquisition could be. Yahoo having the research
knowledge and Microsoft the funds as well as the software ownership.
Regarding Baidu we cannot clearly see how they could compete against
Google outside of China.
As I said previously specialized search engines are limited to the website they
are linked to.
We could then think about new comers who starting from nothing could beat
famous search engines in a small period of time, it could have been the success of
some products such as Cuil launched in summer 2008 which received a lot of
advertisement through the news17. But search engines is a very ungrateful world
where visitors are giving no more than one chance: the product works or it does not.


Illustration 7: Results page of Cuil

This is a point that I discovered very quickly and that you can test by
yourself. People want the information as soon as they can. They are
ready to test the product but in a certain amount of tries. When you
move from Google to another search engine you are often intransigent. At the first
result which does not fit your expectations you will go back to Google. But is the
search engine wrong or is it because it is responding differently that on what you
were used to?
In order to conclude this part I would say that with the search engine history
we have and the search engine market configuration, I cannot see how Google
could lose its position. Until now only one company succeeds to make a such gap in
the world of search engine and it is Google itself and it was in a period where
everything had to be created on Internet.
So I would say that on this field I don't see how Google can be beaten and
even worried.
A new technology regarding research is however more and more recurrent in
this field and is called semantic research.

3.1.6 The semantic web

“This part needs additional information and improvements and is then not finished
yet.”

The semantic web is another way of crawling the net. We all know how to
make a search on the Internet isn't it? We just type in some keywords and press the
return key in order to get the answer. In this configuration you have to feed the
search engine with the request.
With semantic Web the concept is a bit different and based on suggesting you
the request instead of typing it entirely. Each time that you are starting to type your
request a list of suggestions are coming to you. We are recently seeing more and
more this technology on the biggest search engines.
The purpose is in fact to guide you as best as they can in order to put you on
the right track and trying to avoid you to reach the labyrinth of the web.
This technology fits one of the main drawback of search engines and that I
call « search engine technology awareness » which consists in how to write good
requests for search engines.
The main drawback of the semantic web is that this is a very new technology
which then have a lot to do before reaching his maturity point. Here we are speaking
about a maximum length of a decade. We can also complain about the rigidity of the
system but it is true that with length and experience this issue could be fixed.
It said that Ask is one of the search engine which based a lot of R&D on this
new technology but according to me and without being a technician I think that
Google can have better results because of its huge database of requests. Future will
tell us what is going to happen.

3.2 Search engines dependency aspect

As I mentioned it in the introduction I define search engine dependency as the
fact that people are swearing only by one search engine when looking for
information on the Internet.

3.2.1 Search engines dependency proves

“This part needs additional information and improvements and is then not finished
yet.”

This is not the studies which are lacking on this topic. When looking for
information regarding information literacy on the Internet you arrive on different
sources of studies and this regarding all the countries of the world. Information
literacy is a relevant topic and an issue. I focused on some very recent and
francophone research that I found on the Internet regarding Canadian students18,
French19 and Belgium students20. I also found information regarding Germany on this
topic. Many sources are as well saying that such research have been made in mostly
all Europe, China (Hong Kong)21 and the United States22.
All the studies I found until now (all done on students panels so literate
people) are all saying the same thing: search engine are the first source of
information when looking on the Internet and all students seem to have receive not
enough training on how to look for information on the Internet.
The best study I found on this topic is one made on all the registered PhD
students (2,218 with an answer rate of 23,4%) last year (2008) on a whole region of
France (not a high technological developed country but far to be the least on a
worldwide scale)23.
As we can imagine PhD students have a high requirements regarding quality
of information.
The study shows that 67,5% of the respondents have never received a training
regarding how to look for information during their whole stay at the university which
could explain the fact that people are running toward search engines directly.
Search engines are used in 96% of the cases when performing research
(which emphasize the necessity of how to well use those technologies).
94% of them do not use blogs which I take as a good thing (even if the survey
is saying the opposite, blogs being written by professionals as well).
The most used search engines are Google (85%) and Google Scholar 37%
(which is a sub search engine of Google).
60% of them do not know what is a meta search engine and only 5% of them
use them.
46 % do not know the search engine of their field and only 20% do use them.
Those figures are very interesting because they show clearly how people are
not adapted to the technology they are using. PhD students should be some of the
most search engines awarded people and it seems that for France they are not. They
are strictly dependent of a single search engine which is here Google. They know
very few of his sub search engine and as written above they do not know how to use
the technology. They are also not aware massively about other search engines.