Showing posts with label google. Show all posts
Showing posts with label google. Show all posts

Monday, July 28, 2008

Google findet mehr als 1 Billion(!) Web-Adressen...

Ich hatte es ja schon immer gewusst, dass der Gogle Suchindex "ziemlich" gross ist. Die letzten "offiziellen" Zahlenangaben, die mehr oder minder indirekt gemacht wurden, besagten, dass Google im Jahr 2005 einen Datenbestand von 24 Milliarden Webseiten im Index verwaltet [1]. Aber die Zeit bleibt ja nicht stehen und das Web wächst beständig....und jetzt schreibt der "official Google Blog" am Wochenende, das der Google-Suchindex die magische Marke von 1 Billion (!) Webseiten überschritten hätte.....[2]

Natürlich muss man bei US-amerikanischen Zahlenangaben stets vorsichtig sein. "One billion" steht ja lediglich für unsere "Milliarde". "Eine Billion" dagegen sind tatsächlich 10^12 (1.000.000.000.000), im Englischen "one trillion", also eine ganze Menge. Jetzt stellen wir uns einmal vor, wir haben diese Billion Indexeinträge, die zudem noch untereinander verlinkt sind. Würde man diese Datenstruktur klassischerweise als Matrix speichern, bräuchte man 10^24 Einträge, von denen die allermeisten ja leer wären. Also speichert man eine derartige Datenstruktur doch besser auf effizientere Weise. Allerdings muss man dabei bedenken, dass der Zugriff auf Links immer noch sehr schnell erfolgen muss, da die iterative PageRank-Berechnung ja auch nicht ohne ist [3]. Ich wäre wirklich einmal daran interessiert, wie lange jetzt eigentlich eine komplette Berechnung des PageRanks für den Gesamt-Datenbestand heute dauert....

[NACHTRAG:]
Tja...man soll ja den Tag nicht vor dem Abend loben...
Der San Francisco Chronicle setzte heute einen Nachtrag zu o.a. Google Meldung, in der es hieß, dass der GoogleBot zwar mehr als 1 Billion Webseiten gefunden hätte, von diesen aber lediglich 30 - 50 Milliarden im Google-Suchindex verwaltet werden [4]. Naja, immerhin haben wir jetzt einen Anhaltspunkt, wie groß das WWW sein könnte.....und dass tatsächlich auch nicht alles bei Google gefunden werden kann.


References:
[1] TNL Blog: Google has 24 billion items index, considers MSN search nearest competitor, September 2005.
[2] The Official Google Blog: We knew the Web was Big....., Juli 25, 2008.
[3] Sergey Brin and Lawrence Page: The anatomy of a large-scale hypertextual Web search engine, Computer Networks and ISDN Systems30(1-7):107--117(1998).
[4] SFGate: New Search Enging challenges Google, July 28, 2008.

Friday, July 11, 2008

Ceci n'est pas Stockholm.....


...or even more "ceci n'est pas Google?"

Very nice allusion to Google's image search found at blogoscoped.com. I hope we will soon find more of that kind :)

Friday, April 11, 2008

Alive and Kicking... Google 2084

Yep...after several months of abstinence, I'm back, finally.
During the last few months, a lot has happened, in science, in business, and also in private. Thus, there is a lot to talk about and a lot of forthcoming posts. Probably the most complex thing during the last months was to handle two large EU-FP7 proposals as being one of the proposal's core partners, with writing, discussing, conferencing, travelling, conferencing, discussing, and writing again, etc....
Also in my genuine research areas a lot has happened. You might be looking forward to several blog posts on information retrieval, semantic web, semantic search, multimedia retrieval, web of trust, e-Learning, and many more.

Today I just wanted to show you a funny 'screenshot' of what might Google look like in 2084.

Ok. The image is not really new. Actually it was published by Randy Siegel back in 2005. You might also find it in the New York Times. The question is not, if Google would be capable to offer these services, but rather when. Anyway, I guess theese services could be offered much earlier......and (at least) also with a much better user interface.

Wednesday, August 29, 2007

Combining Social Networking and Traditional Web Search

We are already waiting quite a while for Google (or Yahoo) to incorporate social networking technology into their web search services. Personalization of search results (as already being offered -- of course in some limited way) as a very first step towards the right direction has become part of traditional web search. The next step should include not only search results from personal resources but also from resources being available within your personal social network. Seems to be straight forward...?!

First thing, include all resources that I have tagged (you can also distinguish between resources that are yours and resources owned by others that you have tagged). In addition to the traditional hyperlink graph of the web (as it is used by the PageRank algorithm), new (weighted and labelled) arcs have to be inserted connecting your homepage (site/blog or whatever) with all the resources (own and foreign) that have been tagged by you.
Ok, now let's consider that everybody's tag-links are inserted into the hyperlink webgraph. Next thing is to insert (weighted and labelled) arcs from your homepage to the homepages of all your friends (according to your personal social network).
You will end up with a graph that includes (a) traditional hyperlinks, (b) personal tagged + weigthed links to resources, and (c) personal tagged + weigthed links to other users. This composite social webgraph should be sufficient for extending the traditional PageRank algorithm to include social networking information.
Of course (and I'm pretty sure of that and I don't go into details now) a lot of adjustments concerning the weights and the use of tags/labels for indexing have to be considered. But, I think it should be possible...

To some extend, a similar approach has been implemented by lijit. lijit is a personalized search engine that makes use of all your available social networking information. At registration, besides your homepage (or blog ... unfortunately you can not manage several different blogs) you pass over your username for several bookmarking/social networking services as well as the URLs of (a) content provided by you and (b) websites (blogs) of your friends. lijit creates a searchable webgraph covering all the resources and networking information that you have provided. Therefore, searching with lijit comes close to a rather personal variant of searching your very own web-universe.

Social Graph Based Search is also the topic of a video podcast from Scobleizer. Although I can't follow his argument that the social networking companies will kick Google's butt in four years (he states that Google's PageRank algorithm cannot be adapted to include social networking information...at least not in a way scalable for Google's purposes and also not with the current business modell of today's SEOs [=Search Engine Optimizers]), I agree to the way how to include social networking information into traditional web search.

Tuesday, August 28, 2007

New Conspiracies ahead - The Google Masterplan

Did you ever wonder what Google is doing with all the terabytes and petabytes of information they collect....? Almost everybody is using any of the services offered by Google. And they keep track of what you are searching (and doing). Do you think they do it merely for placing some (more or less well suited) text ads??

This short video clip from Özan Halici und Jürgen Mayer presented at the EMERGEANDSEE Festival Berlin 07 is raising the same (more or less rhetorical) questions. See for yourself...

[via 'Die ZEIT']

Tuesday, August 07, 2007

Do Students really appreciate online lecture recordings...??

For another paper I was doing some research on how students are using online lecture recordings. I started with asking Google about studies and evaluations on the topic. But, as I had to find out, there are only some primitive evaluations (with a small number of participating students often limited to a spectial course or a single university) as to the satisfaction of the students -- often with 'overwhelming' results (as for advertising the used lecture recording technology...).

Here you may find a few examples (recent evaluations with more than 100 students seem to be rare...):

  • Maree Gosper, Margot McNeill, Karen Woo et. al.: Web-based Lecture Recording Technologies: Do Students Learn From Them?, in In. Adelaide: Higher Education Research and Development Society of Australasia, 2007.
    (interesting summary of several surveys down under in Australia with Lectopia/iLectures, more than 10.000 students addressed but only about 800 answered...)

  • Marc Krüger: Pädagogische Betrachtungen zu Vortragsaufzeichnungen (eLectures), in i-com, Zeitschrift für interaktive und kooperative Medien, Oldenbourg Wissenschaftsverlag 3 56--60 (2005).
    (first try of a more general study...)


I was not able to find a more general, summarizing report dealing with a variety of scenarios and clustering the results (inluding methodological soundness and decent empirical basis).

So...first thing -- if you know about any study or evaluation (with significant empirical basis), please write a comment with a link to the source (!)
Second -- I know that this will also only be a rather limited and primitive poll -- please fill out the small poll below. Maybe we gain a little bit more insight on student's satisfaction concerning online lecture recordings -- not only restricted to a single course or a single university....

Tuesday, May 15, 2007

Next Generation Search Engines revisited...

As reader of my blog for sure you know that one of the main topics is (semantically enhanced) searching the web. Recently I was rereading Andrei Broder's short paper on 'A taxonomy of web search'[1], where he was also referring to the three generations of web search engines. His article dates back to 2002. Environment and technology in the web are rapidly changing. So, what about this three generations? Do we already need a 'next generation'? And what about the discussion about 'Search 2.0'?...
But first, let's recall the three generations according to Broder:

  • First generation: search engines are useing almost only on-page data such as text and formatting information to compute result ranking (1995-1997, cf. Alta Vista, Excite, etc...).

  • Second generation: search engines are using off-page, web-related data such as link analysis, anchor-texts, and click-through data (1998-..., cf. Google).

  • Third generation: search engines try to blend data from multiple, heterogeneous sources trying to answer 'the need behind the query'. The computed results are customized according to the user's information needs, taking into account the user's personal data background, context, and intention (now? - ...).

Clearly, search engines of the third generation include social networking information, tagging, user feedback, semantic analysis, recommendations, and trustworthyness of information (according to its source).
In read/write web the topic is also addressed as comparison of traditional search technologies with what they call 'Search 2.0' [2]. As usual, I don't like the 2.0-term. What is discussed there, refers to Broder's definition of 'Third generation' and adds nothing significant new to it (besides the marketing term). But, the article is definitely worth reading, because a lot of recent search engines are referenced and discussed (and also because of the interesting discussion that follows). They distinguish between 'Finding' information and 'Discovering' information, while relating the second term to 'Search 2.0'.

Broder distinguishes different sorts of web search queries:

  1. Navigational: intented use is to reach a particular web page (similar to 'known item' search in classical information retrieval). Therefore, navigational queries usually do have only one 'right' result.

  2. Informational:
  3. intended use is to acquire information assumed to be present on one or more web pages (as in classical information retrieval).
  4. Transactional:
  5. intended use is to find a web page, where further transactions (e.g. shopping) will take place.

If we take social bookmarking services, navigational queries can be computed simply by using the user's personomy (i.e. the set of all tags used by a distinct user). If the goal is to find a web page, which has been already accessed in the past, the page might be found quickly, if the user has registered the page within the bookmarking service (which comes to 'Finding' information). But, the query might also be resolved by using other people's tags, if somebody has tagged the page (with objectively descriptive tags).
Social bookmarking services are also usefull for the other two purposes. In addition, if the page is found, the social networking information can be utilized for 'discovering' new, previously unkown, but related (similar) information (which comes to 'Discovering' information). Hotho et al. present an adaptive ranking algorithm (FolkRank) for social bookmarking systems and discuss the problems that arise for tag-based search engines [3].

But, to answer the 'need behind the query' as Broder states in his definition of 'Third generation search engines', further personalization is mandatory. Only, if the search engine is able to find out the context of a query w.r.t. a given user and a given situation (i.e. even the same user might have different information needs in different situations), then it is possible to grasp the actual context of the query, and thus, also the 'need behind the query'...
[to be continued]

References:
[1] Andrei Broder: A taxonomy of web search, SIGIR Forum 36, pp. 3-10, 2002.
[2] Ebrahim Ezzy, Richard MacManos: Search 2.0 vs. Traditional Search, read/write web, June 20th, 2006.
[3] Andreas Hotho, Robert Jäschke, Christoph Schmitz and Gerd Stumme: Information Retrieval in Folksonomies: Search and Ranking, in Proceedings of the 3rd European Semantic Web Conference, pp. 411-426, 2006.

Tuesday, May 01, 2007

veni, vidi, emi ... according to Google


Two weeks ago, Google bought Double Click for an incredible amount of money. But for what reason, you might ask....well, to become the dominant internet advertiser. But that is too simple and too short-sighted, as pointed out last week by the German newspaper 'Die Zeit'. According to an article of Götz Hamann, everybody who is worrying about Google to become No. 1 internet advertiser is as short-sighted as this famous Gaulic chief 2000 years ago complaining about Julius Caesar having defeated the entire Gaul. But, Julius Caesar's primary ambition was not only to conquer Gaul, but to become the leader (cesar) of the Roman world empire. And the Google managers .... they also have something else in mind than just becoming the dominant internet advertiser.
They want to preveil their rules in the global advertising market....against the tv networks and also against the newspapers. In relation to the attention that the newspapers get, they receive way too much advertising. This is because of the brokers distributing advertising for the big corporations. These brokers are one of the reasons, why Google is not able to grow as fast as it possibly could. Thus, Google will continue to buy...and who knows, what company will be next....
Alas, and again we will hear: veni, vidi, emi....

(I came, I saw, and I bought..... freely adapted from H.C. aka Gaius Julius Caesar)

Saturday, March 24, 2007

A short note on Semantic Search Engines


Lars featured an article on Hakia, a 'semantic search engine' in his blog today.
Contrarywise to Google, Hakia uses natural language processing (NLP) to 'understand' search queries given in natural language (and not as plain keywords). Ok, this is now new.
Just remember AskJeeves, a.k.a. Ask.com. But, in difference to that search engine, Hakia claims to perform a 'semantic search' (and not a keyword based search). If you read a little bit further in their technical description at Hakia labs you will find thet they are using a parser called 'OntoSem', which as they claim is able to perform a 'deep semantic analysis' of sentences. This parser is used for query string analysis and also for the generation of the search index. But, their search index is different to that we are used from 'traditional' search engines.
Just remember, traditional keyword based search engines extract so called 'descriptors' from web pages that are used to describe the content of those pages. These descriptors are managed within an inverted index. Thus, by accessing the inverted index with a descriptor (=keyword in query string), a list of web pages will be returned together with some weights indicating the relevance of the descriptor for the page.
So, how does the index work in Hakia? They clain, that their parser is analyzing all web pages to be indexed sentence by sentence. In effect, all possible questions that can be posed for each sentence are generated, forming what they call the 'QDEX data'. All possible questions for all sentences in all pages have to be stored in an index-like data structure (simply for fast and efficient access). If now a query string contains a question, this question is mapped against the index, resulting in a large number of 'relevant' (questions, sentences and in the end...) web pages. Now they apply a 'smart' algorithm called 'Semantic Rank' which orders the resulting list of documents according to their relevance wrt. the question given in the query string. More details about the technique is not published (at least not to my knowledge).

Thus, the only way to find out about the quality of their approach is to try out their search engine. A nice feature is that they have included small example applications where you may try out OntoSem or QDEX interactively by yourself. I have only tried OntoSem (because for QDEX you have to sign up a request) and the result was not really different from any standard NLP-Parser (and thus not really convincing. I will sign up and give QDEX a try..and of course I will post the result).

If I compare their approach to Google's, the problem is that Google is also claiming to incorporate semantic technology altogether. So, again I can only compare the results of my queries. I have tried several queries and have come to the following results:

  • in general the results from hakia do look very promising!

  • in comparison to Google results, they are not really that different this may come because of my 'queries' and thus, my results are probably not really objective).
    Let me give you an example: I was asking both (Google and Hakia) the question 'When did the Semantic Web start?'. Ok, it's not so easy to understand at all. First, it didn't really start (as an being implementation), but the topic itself startet several years ago...and so I was curious about the results:

Ok, as stated before, this is not representative at all, but I will keep an eye on Hakia anyway. As soon as I have more representative experiences (and hopefully some comments or other interesting comparisons), I will write about.

One thing in the end. After all, what I have read about Hakia and on their web site, they seem to have a completely different notion of 'Semantic Search'. What they do is to apply (enhanced) information retrieval techniques based on natural language processing. They do NOT evaluate any additional given semantic annotation (RDF, OWL, etc...), which might be addid into the web pages or connected to them. They claim (same as Peter Norwig) that the average user is much too stupid to supply semantic annotation (because for doing that he needs to be a linguist as they say).
Of course, if you are following the PLAIN trail of the W3C and encode everything by hand (HTML, XML, RDF, OWL, SWRL, etc...) then you have to be some real expert. But today, you (at least most of you) don't encode HTML by your own. Remember, there are lots of real nice WYSIWYG editors for doing that AND (even more) there are several simplifications (just think of writing your blog). In the same way there will be smart user interfaces and editors providing help in generating semantic annotations for your texts. First and most simple step is providing labels and keywords (just 'tags') for your blog posts. Most people do that...just because (1) they want to put some order into the set of their postings and (2) they want their posting to be found.

Friday, March 23, 2007

Who's afraid of Google...?


Today I read an article in the German newspaper 'DIE ZEIT' entitled with 'Who's afraid of Google'. It was about Google's project of digitizing the entire printed books of the world and their deal with the 'Bayerische Staatsbibliothek' (the 2nd largest scientific library in Germany). Several scientists and librarians were asked about what they think about Google's plans and its effects on the culture in general.
In general the opinion was more critical than enthusiastic (and I can follow their arguments). On the one hand, Google offers some kind of democratozation of the reading culture (meaning that in the end all books might be available for everybody at any time). But, Google as a philanthropist...? Business first! This means, clicks are money...and thus, it's all about money. The main criticism of the community was that Google will get a monopoly concerning all our reading. Thus, the obvious next step (from some pessimistic point of view) is censorship (or at least 'filtering'). Just imagine Google to be in a position that dictates what we are reading (I mean concerning the content). Then, there is no democracy at all....
The other complaint of the librarians was concerning the poor quality of Google scanning the books. There are examples of 'thumbs' inside scans (btw I was not able to find an image proof of that...but I confess that I was searching with Google...) and the offered resolution of 300 dpi for color/grayscale (for handwritten codices as e.g. the 'Sachsenspiegel' this is not enough) and 600 dpi for print is also subject of criticism.

Speaking as a scientist, I'm really happy that almost all scientific literature (at least for me as being a computer scientist) is available in the world wide web and that it can be accessed by searching Google. I don't want to stay in libraries for hours, discussing with librarians, filling out forms, copying articles from books, and maybe waiting weeks or months before I am able to access certain literature. But...(and now be honest)...who of you does not print out an article for reading (and commenting, annotating, etc...).
Google print maybe will offer the possibility to read the very first edition of 'Robinson Crusoe' while sunbathing at the beach...but will you? For sure, we all LOVE books. But, do we also 'love' a computer screen? You might touch books, feel (and scent) the old leather of a binding, browsing through worn pages, making annotations and remarks....just having the ability to put content (knowledge) into some matter (the book), take it everywhere you want to, being your companion...and you know where it is (on the shelf, on your nightstand, in the bathroom...). Do you think, having a book just on some screen is something similar (concerning your sensation). You switch the screen off...and the book is gone. Of course, it's somewhere in the digital universe...but not for real on your nightstand when you put out the light.....

In the end, I don't think that Google print will be as 'mean' as being projected by many librarians. Of course it will bring some change (and also some positive change). But it will not 'extinguish' the book. Publishers and booksellers will keep on selling books and we will also keep on reading them. Monopolies are not democratic (...some contradiction ahead). Therefore, we should also support all the other digitizing projects, as e.g. the digitizing project of the 'Börsenverein des deutsche Buchhandels' or the 'European Library' and their 'Europeana' project that will be launched by the end of march.