Salut à tous,
Comme promis je vous publie mon rapport intermédiaire de thèse.
Pour le télécharger cliquez sur le lien suivant (mais il vous faudra un compte gratuit sur slideshares):
Lien pour la thèse
Le rapport final est prévu pour juin.
Bonne lecture.
dimanche 1 février 2009
Risks of search engine dependency and its influence on data quality
Thesis intermediate report submitted for the European Master in Business Studies
(EMBS)
by Ronan CHARDONNEAU
Institut de Management de l'Université de Savoie d'Annecy (FR)
Università degli studi di Trento (IT)
Universität Kassel (GER)
Universidad de León (SP)
Date of submission: 26th January, 2009
Master Thesis
(EMBS)
by Ronan CHARDONNEAU
Institut de Management de l'Université de Savoie d'Annecy (FR)
Università degli studi di Trento (IT)
Universität Kassel (GER)
Universidad de León (SP)
Date of submission: 26th January, 2009
Master Thesis
Contents
Foreword.......................................................................................................................6
Chapter 1: Introduction of the topic background..........................................................8
1.1 Relevance of the subject...................................................................................10
1.2 Major terms......................................................................................................11
1.3 Focus, goals and structure of the report...........................................................11
Chapter 2: Concept of data quality.............................................................................13
2.1 Data quality definition......................................................................................14
2.2 The importance of data quality.........................................................................15
Chapter 3: Search engines dependency.......................................................................16
3.1 Search engine market configuration.................................................................17
3.1.1 Search engine categories..........................................................................17
3.1.2 Search engine market...............................................................................19
3.1.3 The search engines in the world...............................................................19
3.1.4 The search engine market shares per country...........................................22
3.1.5 The search engines competition...............................................................23
3.1.6 The semantic web.....................................................................................24
3.2 Search engines dependency aspect...................................................................25
3.2.1 Search engines dependency proves..........................................................25
3.2.2 Search engines dependency aspect...........................................................27
3.3 Search engines dependency problems..............................................................28
3.3.1 Privacy issues...........................................................................................29
3.3.2 Looking for other search engines.............................................................30
3.3.3 Search engine awareness..........................................................................30
3.3.4 Other search engines existence awareness...............................................32
3.3.5 Less confident regarding other search engines.........................................33
3.3.5 Less confident regarding other search engines.........................................33
3.3.6 Even the best cannot provide you everything...........................................34
Chapter 4: Risks of search engines dependency and its influence on data quality.....35
4.1 The information has been found but is poor....................................................36
4.2 What the search engines do not tell you...........................................................36
4.3 The best way to get data quality.......................................................................37
4.3.1 The sub-search engines.............................................................................37
4.3.2 The size of the Internet.............................................................................38
4.3.3 Single search engine Internet coverage....................................................39
4.3.4 Multiple search engine Internet coverage.................................................42
4.3.5 Others search engine Internet coverage....................................................44
4.3.6 A concrete representation of the World Wide Web...................................46
4.4 The gap between search engine dependency and data quality.........................47
Chapter 5: The Google example.................................................................................50
5.1 Google..............................................................................................................51
5.2 Google's success...............................................................................................51
5.3 Google dependency state..................................................................................52
5.4 Google functions..............................................................................................52
5.5 Google added functionalities............................................................................53
5.6 Google success is his weakness.......................................................................53
5.7 Google's disappearance hypothesis..................................................................54
Conclusion..................................................................................................................55
Declaration..................................................................................................................56
List of literature...........................................................................................................57
Afterword....................................................................................................................61
Chapter 1: Introduction of the topic background..........................................................8
1.1 Relevance of the subject...................................................................................10
1.2 Major terms......................................................................................................11
1.3 Focus, goals and structure of the report...........................................................11
Chapter 2: Concept of data quality.............................................................................13
2.1 Data quality definition......................................................................................14
2.2 The importance of data quality.........................................................................15
Chapter 3: Search engines dependency.......................................................................16
3.1 Search engine market configuration.................................................................17
3.1.1 Search engine categories..........................................................................17
3.1.2 Search engine market...............................................................................19
3.1.3 The search engines in the world...............................................................19
3.1.4 The search engine market shares per country...........................................22
3.1.5 The search engines competition...............................................................23
3.1.6 The semantic web.....................................................................................24
3.2 Search engines dependency aspect...................................................................25
3.2.1 Search engines dependency proves..........................................................25
3.2.2 Search engines dependency aspect...........................................................27
3.3 Search engines dependency problems..............................................................28
3.3.1 Privacy issues...........................................................................................29
3.3.2 Looking for other search engines.............................................................30
3.3.3 Search engine awareness..........................................................................30
3.3.4 Other search engines existence awareness...............................................32
3.3.5 Less confident regarding other search engines.........................................33
3.3.5 Less confident regarding other search engines.........................................33
3.3.6 Even the best cannot provide you everything...........................................34
Chapter 4: Risks of search engines dependency and its influence on data quality.....35
4.1 The information has been found but is poor....................................................36
4.2 What the search engines do not tell you...........................................................36
4.3 The best way to get data quality.......................................................................37
4.3.1 The sub-search engines.............................................................................37
4.3.2 The size of the Internet.............................................................................38
4.3.3 Single search engine Internet coverage....................................................39
4.3.4 Multiple search engine Internet coverage.................................................42
4.3.5 Others search engine Internet coverage....................................................44
4.3.6 A concrete representation of the World Wide Web...................................46
4.4 The gap between search engine dependency and data quality.........................47
Chapter 5: The Google example.................................................................................50
5.1 Google..............................................................................................................51
5.2 Google's success...............................................................................................51
5.3 Google dependency state..................................................................................52
5.4 Google functions..............................................................................................52
5.5 Google added functionalities............................................................................53
5.6 Google success is his weakness.......................................................................53
5.7 Google's disappearance hypothesis..................................................................54
Conclusion..................................................................................................................55
Declaration..................................................................................................................56
List of literature...........................................................................................................57
Afterword....................................................................................................................61
Foreword
As most of the students who has a computer one of my first move when I
wake up is to switch on the computer and to spend my first twenty minutes of the day
on the Internet.
From there I have a look at the last news, I check my e-mails and eventually
exchange some few words with a couple of friends by using online chat applications.
I also check my other email account as well as my blogs and analyze the traffic I got
during the last few days, to finish this process I consult my advertisement account to
see if I got some revenues. I often use as well search engine to look for information
which just came up into my mind during the night.
In the paragraph you just read was the description of my morning routine on
Internet. There is nothing special except that most of the moves I described above are
in fact done on two to three major search engines: Google, Yahoo and Microsoft.
I hardly ever use Yahoo or Microsoft for search purpose but Google is for
sure the website I visit the most to crawl the web but... is Google the Internet?
I got the idea to write about: « Risks of search engine dependency and its
influence on data quality » not because I was using all those Google applications
everyday and was scared about what will happen if I get in troubles with Google
such as privacy issues or if Google just closed. I just write about it because one day I
found Google results not accurate enough.
And from this observation a lot of questions came to my mind:
· Is it me who is not good enough at performing research on the Internet?
· Is it because no one wrote about the information I am looking for?
· Is it because the information is not on the first pages in Google that I have to
browse all the pages in order to find it?
· Is it because Google is not good enough?
· Is it because the information is hidden in some other documents such as PDF,
pictures, videos?
· Is it because I have to use another way to crawl the web and if yes how?
You see here how a simple observation can raise a lot of questions.
I hesitated a lot about writing on this topic, the main problem I got was that I
was not convinced that there is a potential risk of being search engine dependent. The
reason is that companies such as Google are working hard in order to fit Internet
users expectations and the vision we get is that they are doing a wonderful work. The
problem is that there could be a difference between perception and real facts and this
is exactly what I am eager to discover here.
Can we measure how huge is the gap between the information we were
looking for and the one of search engines as Google are providing us?
Search engines are set up to find information on the Internet, information
being the basis of any good decisions making we can then understand how important
and interesting it is to write on this topic.
I hope you will appreciate this reading as much as I did when making my
research.
wake up is to switch on the computer and to spend my first twenty minutes of the day
on the Internet.
From there I have a look at the last news, I check my e-mails and eventually
exchange some few words with a couple of friends by using online chat applications.
I also check my other email account as well as my blogs and analyze the traffic I got
during the last few days, to finish this process I consult my advertisement account to
see if I got some revenues. I often use as well search engine to look for information
which just came up into my mind during the night.
In the paragraph you just read was the description of my morning routine on
Internet. There is nothing special except that most of the moves I described above are
in fact done on two to three major search engines: Google, Yahoo and Microsoft.
I hardly ever use Yahoo or Microsoft for search purpose but Google is for
sure the website I visit the most to crawl the web but... is Google the Internet?
I got the idea to write about: « Risks of search engine dependency and its
influence on data quality » not because I was using all those Google applications
everyday and was scared about what will happen if I get in troubles with Google
such as privacy issues or if Google just closed. I just write about it because one day I
found Google results not accurate enough.
And from this observation a lot of questions came to my mind:
· Is it me who is not good enough at performing research on the Internet?
· Is it because no one wrote about the information I am looking for?
· Is it because the information is not on the first pages in Google that I have to
browse all the pages in order to find it?
· Is it because Google is not good enough?
· Is it because the information is hidden in some other documents such as PDF,
pictures, videos?
· Is it because I have to use another way to crawl the web and if yes how?
You see here how a simple observation can raise a lot of questions.
I hesitated a lot about writing on this topic, the main problem I got was that I
was not convinced that there is a potential risk of being search engine dependent. The
reason is that companies such as Google are working hard in order to fit Internet
users expectations and the vision we get is that they are doing a wonderful work. The
problem is that there could be a difference between perception and real facts and this
is exactly what I am eager to discover here.
Can we measure how huge is the gap between the information we were
looking for and the one of search engines as Google are providing us?
Search engines are set up to find information on the Internet, information
being the basis of any good decisions making we can then understand how important
and interesting it is to write on this topic.
I hope you will appreciate this reading as much as I did when making my
research.
Chapter 1: Introduction of the topic background
I will not surprise you if I say that Internet has been created to share
information and to communicate with each others.
It is hard to evaluate how big is the Internet, estimations among companies
are very different, it varies from 15 to some 30 billion Web pages1. The number of
websites is increasing everyday and estimated at 185,167,8972 with a constant
augmentation since the creation of the world wide web.

diverse manners if it comes to be a standard for companies (recognized as a mark of
trust, seriousness and quality) it is also a space for many individuals (blog
phenomenon). As an example regarding France, in June 2008 14% of French people
above 12 year-old which means 22% of French Internet users are authors of a blog or
a website3.
The banalization of the Internet and the fact that anyone can create his own
website for free increase the feeling we have regarding the Internet: a true jungle of
information and even sometimes real “dump” regarding information accuracy.
Websites can be accessible through three channels:
· Direct access (for example you know the website address by heart, you put it
in your favorites or you find a website on a business card and you are typing
it in the address bar);
· External links (you access to a website which has the link of another
website, this is the case in most of websites, catalogs, advertisement);
· Through Search Engines (you use a dedicated application by typing in some
keywords in order to get suggestions of what you are looking for);
As you can see from this list if you use only the first two ways to crawl the
web it comes to be too rigid and not wide enough. It has been said as well that the
first way is disappearing more and more in profit of search engines4.
So one could say that there is currently two main ways to crawl the web, from
link to link and by using search engine.
This last one being indispensable in order to crawl the web properly.
More and more information are put on the Internet which makes it
come a true jungle. The only way to crawl those information properly
is to use search engines.
information and to communicate with each others.
It is hard to evaluate how big is the Internet, estimations among companies
are very different, it varies from 15 to some 30 billion Web pages1. The number of
websites is increasing everyday and estimated at 185,167,8972 with a constant
augmentation since the creation of the world wide web.

Illustration 1: Total Sites Across All Domains August 1995 - January 2009
Habits have changed since the creation of the Internet and websites are used now indiverse manners if it comes to be a standard for companies (recognized as a mark of
trust, seriousness and quality) it is also a space for many individuals (blog
phenomenon). As an example regarding France, in June 2008 14% of French people
above 12 year-old which means 22% of French Internet users are authors of a blog or
a website3.
The banalization of the Internet and the fact that anyone can create his own
website for free increase the feeling we have regarding the Internet: a true jungle of
information and even sometimes real “dump” regarding information accuracy.
Websites can be accessible through three channels:
· Direct access (for example you know the website address by heart, you put it
in your favorites or you find a website on a business card and you are typing
it in the address bar);
· External links (you access to a website which has the link of another
website, this is the case in most of websites, catalogs, advertisement);
· Through Search Engines (you use a dedicated application by typing in some
keywords in order to get suggestions of what you are looking for);
As you can see from this list if you use only the first two ways to crawl the
web it comes to be too rigid and not wide enough. It has been said as well that the
first way is disappearing more and more in profit of search engines4.
So one could say that there is currently two main ways to crawl the web, from
link to link and by using search engine.
This last one being indispensable in order to crawl the web properly.
More and more information are put on the Internet which makes it
come a true jungle. The only way to crawl those information properly
is to use search engines.
1.1 Relevance of the subject
Internet is becoming more and more our information provider. "In 2002 a
study from the U.S. Department of Commerce estimated that 36% of the American
public over the age of 3 used the Internet to search for product and service
information in 2001. This usage represented a substantial increase in information
search behavior from 2000, when 26% of the American public reported using the
Internet to search for product and service information."5
Between 2000 and 2008 the USA got an increase of Internet users of 130,9
%6 we can then imagine how important is searching information through Internet
nowadays.
The number of Internet users is estimated to 1,463,632,361 (world population
6,676,120,288) with a growth rate from 2000-2008 fixed at 305.5 %7.
study from the U.S. Department of Commerce estimated that 36% of the American
public over the age of 3 used the Internet to search for product and service
information in 2001. This usage represented a substantial increase in information
search behavior from 2000, when 26% of the American public reported using the
Internet to search for product and service information."5
Between 2000 and 2008 the USA got an increase of Internet users of 130,9
%6 we can then imagine how important is searching information through Internet
nowadays.
The number of Internet users is estimated to 1,463,632,361 (world population
6,676,120,288) with a growth rate from 2000-2008 fixed at 305.5 %7.
1.2 Major terms
In this thesis you will hear a lot about the following terms: search engines,
search engine dependency and data quality.
Search engine is the most flexible technology which has been created in
order to crawl the web. A search engine is no more than a web application which is
processing data. A search engine does not create data it just process some
information it has in his index.
I describe search engine dependency as the fact that one does not make the
choice between one search engine and another. Search engine dependency is very
relevant in most of the countries. Most of the people are using only one single search
engine when performing requests on the web. I will speak as well about bad habits
when dealing with search engines. For example you can be dependent of using a
search engine but using it badly.
Data quality is the quality of data. Data are of high quality "if they are fit for
their intended uses in operations, decision making and planning ". Alternatively, the
data are deemed of high quality if they correctly represent the real-world construct to
which they refer. These two views can often be in disagreement, even about the same
set of data used for the same purpose8.
search engine dependency and data quality.
Search engine is the most flexible technology which has been created in
order to crawl the web. A search engine is no more than a web application which is
processing data. A search engine does not create data it just process some
information it has in his index.
I describe search engine dependency as the fact that one does not make the
choice between one search engine and another. Search engine dependency is very
relevant in most of the countries. Most of the people are using only one single search
engine when performing requests on the web. I will speak as well about bad habits
when dealing with search engines. For example you can be dependent of using a
search engine but using it badly.
Data quality is the quality of data. Data are of high quality "if they are fit for
their intended uses in operations, decision making and planning ". Alternatively, the
data are deemed of high quality if they correctly represent the real-world construct to
which they refer. These two views can often be in disagreement, even about the same
set of data used for the same purpose8.
Inscription à :
Articles (Atom)