Tuesday, March 27, 2007

Online Umfrage zum IT Gipfel 2006 ....


Mein lieber Kollege Justus Bross möchte eine online Umfrage bzgl. der Einrichtung eines Blog zum IT Gipfel der Bundesregierung (hier auch ein Video-Podcast zum Thema) im vergangenen Jahr vornehmen. Ziel ist dabei die Einrichtung eines 'Gipfelblogs', das -- da unter der Flagge der Bundesregierung -- spezielle Vorkehrungen in Sachen Moderation und redaktionelle Kontrolle beinhalten soll (und über die wir derzeit noch diskutieren).

Daher also hier der Aufruf, sich kurz 10 Minuten Zeit zu nehmen und schnell über den Fragebogen zu gehen. Ihr würdet uns damit sehr helfen.

Hier der Link zum online-Fragebogen
....und für eine Weiterverbreitung wären wir natürlich ebenso sehr dankbar.

...und so geht unser Dank bereits an

ACHTUNG: Besten Dank an alle Teilnehmer und Teilnhmerinnen!!
Wir haben knapp 200 ausgefüllte Fragebögen zurückerhalten und die Aktion beendet.
Sobald die Auswertung abgeschlossen ist, werden wir eine erste Bilanz ziehen.

From iRack to iRan.....

I guess apple never thought (but maybe dreamed) of extending their iProduct line into domestic appliances or even sports ... :)


Link: sevenload.com

Monday, March 26, 2007

Alexandre Dumas - Die drei Musketiere (The three Musketeers)


For sure you will think 'hey, why is he reading that old children's book again?', but no, you're wrong. I've read this masterpiece of Alexandre Dumas for the very first time. Of course, everybody knows the story. We all have seen so many movies and remakes of movies, and the musketeers have become a cliché even long ago. But, I had the chance to get a rather old copy. It's an almost 100 years old german translation (of course in old german typesetting). And one thing about the book is, it's language is remarkable and so different compared to today's german. But before I'll go into details, let me rather briefely recapitulate the story (in german....):

Wir schreiben das Jahr 1625. D'Artagnan, der arme Sohn eines Landadeligen aus der Gascogne will sein Glück als Musketier in der Garde des Herrn von Treville - also der Leibgarde des Königs Ludwig XIII. - versuchen. Ausgestattet mit einem Empfehlungsschreiben seines Vaters, einem altgedienten Musketier gerät er, der sich wie alle Gasconier (so Dumas) leicht in seiner Ehre gekränkt sieht, in eine Auseinandersetzung mit einem mysteriösen Gardisten des Kardinals Richelieu. Bevor es aber zu einem Duell kommt, wird er vom Küchenpersonal der Landschenke, in der er Rast macht, niedergeschlagen und seines Empfehlungsschreibens beraubt.
In Paris angekommen versucht er sich dennoch um eine Audienz bei Herrn von Treville und verstrickt sich im Laufe des Tages - natürlich bedingt durch sein aufbrausendes Temperament ebenso wie durch seine eigene Tollpatschigkeit - in sage und schreibe drei Duelle mit drei Musketieren des Königs, namentlich mit den drei Freunden Athos, Aramis und Portos. Verabredet im Park von Saint Germaine werden alle vier überrascht von einer Patrouille des Kardinals. D'Artagnan bietet den drei Freunden seine Hilfe beim bevorstehenden Degengefecht an, aus dem die Musketiere siegreich hervorgehen.....und das ist der Beginn einer Reihe äußerst spannend zu lesender Abenteuer.....
Die eigentliche Handlung ist ja aus diversen Filmen bekannt, allerding ist die Charakterzeichnung Dumas sprachlich natürlich noch wesentlich reizvoller als deren filmische Umsetzung. Für mich wird wohl immer die Verfilmung aus den 70er Jahren mit Michael York als D'Artagnan, Oliver Reed als Athos und Richard Chamberlain als Aramis (wer den Portos gespielt hat fällt mir nicht ein....war aber jemand, den ich sonst nicht weiter kannte) DIE VERFILMUNG überhaupt bleiben. Besonders schön - neben einem beeindruckenden Oliver Reed und unberechenbaren Charlton Heston als Kardinal Richelieu - war wohl die Tatsache, dass sich anscheinend keine der Romanfiguren (außer den Bösewichtern) in dieser Verfilmung wirklich ernst nimmt. Ach und natürlich Racquel Welch als wohlgefomte und unwiderstehliche Constanze, Geraldine Chaplin als porzellanfarbene, zerbrechliche Königin und Faye Dunaway als durchtrieben böse verführerische Mylady deWinter. Der Film (d.h. die Verfilmung des kompletten Romans) kam damals ja in zwei Teilen in die Kinos und es braucht seine Zeit, die über 700 Seiten des Romans in Szene zu setzen. Ach, fast hätte ich den Erzbösewicht Rouchford vergessen, den Mann mit der Narbe, hinter dem D'Artagnan den ganzen Roman über her ist. Der wird in besagter Verfilmung von Christopher Lee (unser aller 60er/70er Jahre Dracula) gespielt.
Daneben gabe es zahlreiche weitere Verfilmungen an die ich mich erinnern kann. Angefangen von einer UFA Verfilmung aus den 30er Jahren über einen ersten Farbfilm mit Gene Kelly in der Hauptrolle [to be continued.....] und dann natürlich auch noch diverse Verfilmungen aus den 80er Jahren. Die letzte bemerkenswerte Verfilmung des Musketierstoffes betrifft einen späteren Roman Dumas, nämlich 'Den Mann mit der eisernen Maske' (auch unter dem Titel '10 Jahre danach' oder 'Der Vicomte von Bragelonne' erschienen), der vorallem mit seiner Besetzungsliste der bereits etwas gealterten Haudegen brillierte (Gerard Depardieu, Jeremy Irons, John Malkovich....).

Mein Fazit: Der Roman hat es in sich! Kein Wunder, das er sich über 150 Jahre internationaler Beliebtheit erfreuen kann. Natürlich darf man keine allzu tiefgründigen Charakterstudien und innere Monologe erwarten, aber Dumas gelingt es, seine Figuren nicht nur in haarsträubend flott erzählte Abenteuer zu verwickeln, sondern ihnen auch noch ziemlich viel Persönlichkeit mit auf den Weg zu geben. Besonders deutlich wird das an der unberechenbaren Gestalt der Lady deWinter und D'artagnans ambiguen Gefühlsaufwallungen ihr gegenüber. So liebt er sie auf der einen Seite (oder ist zumindest immens verliebt..) wie er sie auf der anderen Seite als Geschöpf des Kardinals zutiefst hasst.
Also, LESEN! und am besten nicht die aktuellen Übersetzungen, sondern eine ältere (vor 1930), um sich an der mitunter barocken Sprachgewalt zu erfreuen.......

Saturday, March 24, 2007

A short note on Semantic Search Engines


Lars featured an article on Hakia, a 'semantic search engine' in his blog today.
Contrarywise to Google, Hakia uses natural language processing (NLP) to 'understand' search queries given in natural language (and not as plain keywords). Ok, this is now new.
Just remember AskJeeves, a.k.a. Ask.com. But, in difference to that search engine, Hakia claims to perform a 'semantic search' (and not a keyword based search). If you read a little bit further in their technical description at Hakia labs you will find thet they are using a parser called 'OntoSem', which as they claim is able to perform a 'deep semantic analysis' of sentences. This parser is used for query string analysis and also for the generation of the search index. But, their search index is different to that we are used from 'traditional' search engines.
Just remember, traditional keyword based search engines extract so called 'descriptors' from web pages that are used to describe the content of those pages. These descriptors are managed within an inverted index. Thus, by accessing the inverted index with a descriptor (=keyword in query string), a list of web pages will be returned together with some weights indicating the relevance of the descriptor for the page.
So, how does the index work in Hakia? They clain, that their parser is analyzing all web pages to be indexed sentence by sentence. In effect, all possible questions that can be posed for each sentence are generated, forming what they call the 'QDEX data'. All possible questions for all sentences in all pages have to be stored in an index-like data structure (simply for fast and efficient access). If now a query string contains a question, this question is mapped against the index, resulting in a large number of 'relevant' (questions, sentences and in the end...) web pages. Now they apply a 'smart' algorithm called 'Semantic Rank' which orders the resulting list of documents according to their relevance wrt. the question given in the query string. More details about the technique is not published (at least not to my knowledge).

Thus, the only way to find out about the quality of their approach is to try out their search engine. A nice feature is that they have included small example applications where you may try out OntoSem or QDEX interactively by yourself. I have only tried OntoSem (because for QDEX you have to sign up a request) and the result was not really different from any standard NLP-Parser (and thus not really convincing. I will sign up and give QDEX a try..and of course I will post the result).

If I compare their approach to Google's, the problem is that Google is also claiming to incorporate semantic technology altogether. So, again I can only compare the results of my queries. I have tried several queries and have come to the following results:

  • in general the results from hakia do look very promising!

  • in comparison to Google results, they are not really that different this may come because of my 'queries' and thus, my results are probably not really objective).
    Let me give you an example: I was asking both (Google and Hakia) the question 'When did the Semantic Web start?'. Ok, it's not so easy to understand at all. First, it didn't really start (as an being implementation), but the topic itself startet several years ago...and so I was curious about the results:

Ok, as stated before, this is not representative at all, but I will keep an eye on Hakia anyway. As soon as I have more representative experiences (and hopefully some comments or other interesting comparisons), I will write about.

One thing in the end. After all, what I have read about Hakia and on their web site, they seem to have a completely different notion of 'Semantic Search'. What they do is to apply (enhanced) information retrieval techniques based on natural language processing. They do NOT evaluate any additional given semantic annotation (RDF, OWL, etc...), which might be addid into the web pages or connected to them. They claim (same as Peter Norwig) that the average user is much too stupid to supply semantic annotation (because for doing that he needs to be a linguist as they say).
Of course, if you are following the PLAIN trail of the W3C and encode everything by hand (HTML, XML, RDF, OWL, SWRL, etc...) then you have to be some real expert. But today, you (at least most of you) don't encode HTML by your own. Remember, there are lots of real nice WYSIWYG editors for doing that AND (even more) there are several simplifications (just think of writing your blog). In the same way there will be smart user interfaces and editors providing help in generating semantic annotations for your texts. First and most simple step is providing labels and keywords (just 'tags') for your blog posts. Most people do that...just because (1) they want to put some order into the set of their postings and (2) they want their posting to be found.

Friday, March 23, 2007

Who's afraid of Google...?


Today I read an article in the German newspaper 'DIE ZEIT' entitled with 'Who's afraid of Google'. It was about Google's project of digitizing the entire printed books of the world and their deal with the 'Bayerische Staatsbibliothek' (the 2nd largest scientific library in Germany). Several scientists and librarians were asked about what they think about Google's plans and its effects on the culture in general.
In general the opinion was more critical than enthusiastic (and I can follow their arguments). On the one hand, Google offers some kind of democratozation of the reading culture (meaning that in the end all books might be available for everybody at any time). But, Google as a philanthropist...? Business first! This means, clicks are money...and thus, it's all about money. The main criticism of the community was that Google will get a monopoly concerning all our reading. Thus, the obvious next step (from some pessimistic point of view) is censorship (or at least 'filtering'). Just imagine Google to be in a position that dictates what we are reading (I mean concerning the content). Then, there is no democracy at all....
The other complaint of the librarians was concerning the poor quality of Google scanning the books. There are examples of 'thumbs' inside scans (btw I was not able to find an image proof of that...but I confess that I was searching with Google...) and the offered resolution of 300 dpi for color/grayscale (for handwritten codices as e.g. the 'Sachsenspiegel' this is not enough) and 600 dpi for print is also subject of criticism.

Speaking as a scientist, I'm really happy that almost all scientific literature (at least for me as being a computer scientist) is available in the world wide web and that it can be accessed by searching Google. I don't want to stay in libraries for hours, discussing with librarians, filling out forms, copying articles from books, and maybe waiting weeks or months before I am able to access certain literature. But...(and now be honest)...who of you does not print out an article for reading (and commenting, annotating, etc...).
Google print maybe will offer the possibility to read the very first edition of 'Robinson Crusoe' while sunbathing at the beach...but will you? For sure, we all LOVE books. But, do we also 'love' a computer screen? You might touch books, feel (and scent) the old leather of a binding, browsing through worn pages, making annotations and remarks....just having the ability to put content (knowledge) into some matter (the book), take it everywhere you want to, being your companion...and you know where it is (on the shelf, on your nightstand, in the bathroom...). Do you think, having a book just on some screen is something similar (concerning your sensation). You switch the screen off...and the book is gone. Of course, it's somewhere in the digital universe...but not for real on your nightstand when you put out the light.....

In the end, I don't think that Google print will be as 'mean' as being projected by many librarians. Of course it will bring some change (and also some positive change). But it will not 'extinguish' the book. Publishers and booksellers will keep on selling books and we will also keep on reading them. Monopolies are not democratic (...some contradiction ahead). Therefore, we should also support all the other digitizing projects, as e.g. the digitizing project of the 'Börsenverein des deutsche Buchhandels' or the 'European Library' and their 'Europeana' project that will be launched by the end of march.

Thursday, March 22, 2007

The Semantic Web will NOT fail .....


...well, maybe it will not develop as projected by the W3C...but there is something evolving :)
Just to answer the question raised at the DERI blog "The Semantic Web will fail?" and to make a statement to Stephen Downes' post on "Why the Semantic Web will fail".
First thing, I wonder why the DERI blog leaves Downes' statement uncommented. Isn't DERI doing a lot of Semantic Web research? Ok, but let's go into detail. Downes' first thesis is

"The Semantic Web will never work because it depends on businesses working together, on them cooperating."

Does it? Why do inter business relations work...? Although there's competition going on, the benefit of sharing common standards is pretty obvious. Maybe you're right that the Semantic Web will not develop as the W3C projects (just remember ISO/OSI and TCP/IP for the sake of standardization). Maybe it's not possible to design the Semantic Web in a top-down manner. But, on the other hand maybe we will see some bottom-up evolution as in the way of the social web (aka web 2.0). And why not thinking of some hybrid aproach? Then, business will adapt sooner or later...and they will. At least, when they face the situation that semantic technology is required to run money making applications. Maybe we are still lightyears away from what is promised by the Semantic Web, but from my point of view we are just about to see some major changes and developments in the upcoming 5-10 years.

Further he says that
"But the big problem is they believed everyone would work together:
- would agree on web standards (hah!)
- would adopt a common vocabulary (you don't say)
- would reliably expose their APIs so anyone could use them (as if)
"

The point is that one of the main purposes of the Semantic Web is to get together heterogeneous data. That means data that are coded with different vocabularies (but containing some additional encoded semantic that states how one vocabulary relates to the terms of another). And just look at RDF. There are already a lot of companies that suport RDF in their applications.

Anyway, the criticism is not new. Just remember Peter Norwig's argument...The nice thing about Stephen Downes' post is the discussion below that is definitely worth while reading...

Tuesday, March 20, 2007

Liquid Browsing - a new paradigm for accessing huge collections of data ?!


As promissed yesterday, I've installed iverse's liquidfile, a kind of substitute (or supplement) for the mac finder. (For all those mac illiterates: the finder is some kind of file system browser). Liquidfile applies the principle of liquid browsing to the filesystem of your computer. To get a short overview, how liquid browsing works, just take a look at the following short movie presentations [1] [2]. It's rather difficult to describe the way how it works with words, because its a visual way of browsing...and thus it is better explained in a visual way also. It's a 2 dimensional visualization, where e.g. the x-axis represents a timeline (as e.g. file creation date) and the y-axis represents all files in an alphabetical order. Then, every file is put on that grid and denoted with a bubble. The size of the bubble (which itself is semi-transparent) reflects the size of the file. The nice thing in general about liquid browsing is that you are able to visualize huge amount of data also on small displays. In that case the mouse pointer acts as some kind of magnifying glass.

As for the file system on your computer, this kind of visualizations has some advantages. First, you can dissolve the entire directory structure (if you want) and look at all files at once (on my computer this were more than 15.000 files). You have the possibility to combine complex (realtime) filtering with several selection mechanisms to find the files you are looking for. Simultaneously, the order in which the files are presented in general stays always the same and does not change. Thus, your visual memory will always be able to memorize the approximate position of a distinct file that you are looking for.
Of course the system has several drawbacks (at least now). My macBook is rather new (january 2007). Thus, almost all of my files have the same creation date (most of them were simply copied to the macBook on the very same day....). Therefore, the time axis will work best, if you start to work with your computer and they will be created/changed over the time.
Another drawback is the limitation of the axes to represent only filenames, filesizes, or creation/change dates. I would like to order files also according to other (also content based) criterias to get a better overview.

But, it's a nice appetizer anyway. What I would like to try out is to visualize relationships between entities (files/documents) with this technique. Just imagine bubbles of different entities and if I focus on one, all other bubbles that represent entities, which are in a certain relationship to the entity being in focus will highlight or move closer (of course in realtime...).

Monday, March 19, 2007

CeBIT Day 3 ... finally


Finally I'm back home at the HPI in Potsdam. Yesterday (Sunday) was my last day at this year's CeBIT computer fair and there were a lot of interesting things going on...
First, by chance I met HBS-fellow Jan Schmidt. Together with Oliver Gassner and Reinhard Karger as host he was talking about the use of networking platforms such as second life, their use, the hype, and social spam (read more about it in Jan's blog).
Another very interesting talk was scheduled afterwards: Carsten Waldeck, founder of iverse was talking about liquidfile - a file browsing application based on liquid browsing. Liquid browsing offers the possibility to arrange and manage a huge number of data items (e.g. documents) in a rather efficient way by a smart 2d-visualization. This principle is now applied to file system management on your local computer. There is an implementation for the mac and I will test it asap to give you a report.
In addition Carsten Waldek and host Reinhard Karger were talking about the plan of installing a social networking platform for archiving intellectual content (not in the sense of intelectual property). But, you should be able to express an idea, to upload your idea in any form of document, and to receive an official timestamp. Of course maintaining a database of 'ideas' will be rather difficult. Thus, all those ideas have to be translated into 'logic' (means with the help of ontology and rule representation languages). Then, if you have a new idea, you can check, if anyone else has already thought about it and if there are some people with similar interest. The other way around, you will be able justify that you have a claim in some idea because you have registered it previously. The trouble is....translate a natural language document into logic representation...you know that this - if done completely automatically - will fail (at least too often....). Thus, we will have to wait, until the semantic web will come.
On the other hand....we have been thinking and thinking again about possible killer applications for the semantic web. This(!) could be one....

BTW, I also met another HBS-fellow Andreas Scheper after the talk (read more about the CeBIT future talks and see some videos in Andreas sein Blog). We decided to meet again at re:publica in berlin, where als the HBS has a so called plugin that will be filled with some fruitful discussion...

So, CeBIT is over for me now... If you want to stay informed about ongoing CeBIT work and CeBIT after hours, keep on reading Miss Marple's Blog.

It's CeBIT time again ... Day 2


Incredible, but true...yesterday evening we've succeeded in walking through ALL CeBIT halls. If you've ever been there, you know that this means something. Somebody told us, that taking the whole 9 yards means walking almost 10 kilometers. Ok, in the evening without the crowded passages its much easier (except the crowded parties of course). And there the CeBIT often offers Live music of exceptional quality.

But before that you stand there at your booth from 9 am to 6 pm, talking to peoples, explaining and demonstrating your project, making new contacts, taking and planning appointments...
Saturday and Sunday are the days of the so called 'Beutelratten' (meaning non-business visitors that are hunting all possible giveaways and gadgets, gathering all in free giveaway bags that they have to carry around everywhere and everytime). But, anyway besides of those, today we also had some rather interesting business visitors.
Tomorrow on Sunday will be my last CeBIT day...and I must say it's okay. Another three days would be really too stressful...

Please don't forget to register at OSOTIS, our new video search engine, and win an iPod!!
(to be continued...)

Saturday, March 17, 2007

It's CeBIT time again ... Day 1

Friday, my first day on CeBIT computer fair. This year, we (Osotis) have a booth at the joint exhibition of 'Mitteldeutschland' in hall 09 - the hall of research and future. Next to us are the other EXIST-SEED funded projects of the FSU Jena -- MobiSoft (watch out for Steffen's Blog, esp. for his pictures) and Navimatrix. On the other hand - my second employer is also present in hall 09 - the Hasso-Plattner-Institute of the University of Potsdam at the joint exhibition of 'Berlin-Brandenburg' with the projects Lock-Keeper (a hardware firewalling system), TeleTask (a simple tele lecturing system), and Tele Lab (a laboratory for security techniques). I will write about the most interesting projects of hall 09 in the upcoming posts.

Yesterday, we had a lot of trouble to keep our system running. But today, Jörg (fortunately) managed to solve the major problems. Thus, I tried to upload a few more videos (including a lot of Berkeley public lectures). If you read this blog, please register at Osotis, because we need a lot of users to keep the system working!

As usual, CeBIT would be half as interesting without it's after business parties. Almost in each of the more than 20 halls there were several parties going on (many including live music, all including free drinks, some including free food...). It's the most difficult task during the day to find out, where and what is going on after business, where to go for 'dinner', for some drinks, and for to listening the best music.

Thursday, March 15, 2007

OSOTIS ...winning an iPod...and the CeBIT rumble starts again


I have already talked about the video search engine OSOTIS, but it has again improved over the time. First at all, what does 'OSOTIS' mean? No, it's not some sort of ancient egyptian god. It's just derived from the botanical name for 'forget-me-not', which is greek 'Myosotis'. So, the name already gives some hint for the offered service:
(1) OSOTIS offers search within videos
(2) right now, most videos available at OSOTIS are academic lecture recordings, ranging from short viseo sequences from the famous Solvay conference in 1927 (where Einstein replied to Bohr that God does not throw dice...) up to lectures from Berkeley, MIT, Stanford, Oxford, or also my lectures at the Friedrich-Schiller-University in Jena (Germany).
(3) OSOTIS does not host the videos (as youTube or Google does). They only provide links to your resources. Nevertheless, OSOTIS downloads the offered video stream for post processing and for generating timed annotations for the video serch.
(4) You can register at OSOTIS (btw if you register before April 15th you have the chance to win an iPod 30GB) and maintain your own video collections, maintain an own user profile, make friends, choose your favourite videos, and (!) you can tag videos.
(5) You can even tag inside video streams. This means that the tagging information also includes time information and that the search is able to replay the video exactly from the right position.
(6) OSOTIS is a social networking tool.
And OSOTIS is at the CeBIT computer fair that has just opened its gates. Visit us at hall 9, D04!
Yes...and tomorrow I will be at CeBIT in Hannover for the next three days. So just stay tuned, because I will write about everything interesting that comes into my way.

Wednesday, March 14, 2007

exams...it's always (most times) the same

Again, today are oral examinations to do. During the semester break, it's the time for most of the exams for the computer science diplomas, masters, or bachelors. I teach several courses that are relevant for students of computer science, as 'web technologies' or 'technical foundations of the internet (TFI)' (and of course 'semantic web')...
Most students come to get examines in 'webtechnologies' or 'TFI', and because these courses are given for senior students, we do oral examinations because there are to few students to justify the labour of preparing a written exam.
But, the point is...most student simply think that web technology or internet technology is something they know already (maybe because it's part of our everyday life) and they don't prepare well or they prepare in the wrong way. I always tell them 'you don't have to give me technical details and parameters. Think global! You have to understand HOW everything is working together. Don't take things granted..ASK if you don't understand or if you cannot answer WHY things are done the way they were...'
It's almost useless. Often I have the feeling that I could talk to the wall instead of human beings capable of listening and understanding. Of course, the facts are quite simple. Everybody has heard of HTTP. Everybody knows that HTTP is used for transporting messages between browser and WWW-server...but, if it comes to (web)-caching, cache control, and stuff like that..and that it is even related to HTTP...I guess most people take that for granted and forget it immediately after having heard.
The same with web-programming. Of course everybody has heard of distributed programming, stubs, skeleton, and so on. But, if you ask, what exactely a stub has to do, to get a procedure call transfered to another computer....
I could tell you hundreds of examples including search engine technology, semantic web, ontologies, web-caching, dynamic HTML, client-side/server-side programming, ...

Sometimes, I'm a little bit frustrated. The students even have the possibility to watch the lectures again from our video database. They even have the possibility to get them all on DVD to watch them again at home. In addition we prepare a lot of material....and what is the outcome? About 2/3 (two thirds) of all exams are dissapointing. Don't get me wrong. Not all of these students do fail the exam. They were prepared an knew quite something...but if you try to figure out how things interact and why certain things are arranged in the way they are....you realize that the students take those things for granted and don't care to take a look behind..

O.k....enough of blustering around. The next exam will be at 3 p.m. Let's hope that sombody is reading this blog :)

Wednesday, March 07, 2007

Daniel Kehlmann - Die Vermessung der Welt


I have no idea whether this book has been already translated in English. If so, read it! (...it has. I've just looked it up at amazon..the title is Measuring the World)
It's some kind of biography of two of the most outstanding people of the 19th century, Carl Friedrich Gauss and Alexander von Humboldt. But beware...it's some very special kind of biography......(By the way, I've visited the 'Alexander-von-Humboldt Gymnasium in Schweinfurt' and made my Abitur there).
I'll switch to German.
Das Buch war für mich wirklich eine der großen Überraschungen in den Neuerscheinungen des letzten Jahres. Der Gattungsbegriff Biografie trifft es nicht ganz, aber irgendwie natürlich doch. Wir stürzen mitten in die Geschichte, als der bereits betagte Gauss nach Berlin aufbrechen muss, um auf Einladung von Humboldt an einem Kongress teilzunehmen. Eigentlich will er ja gar nicht, vorallem nicht aus dem Bett heraus, aus seinem Haus, aus Göttingen...und überhaupt. Köstlich...vor allem auch der Dialog (eigentlich eher ein Monolog) mit Eugen, seinem seiner Ansicht nach 'missratenem' Sohn. Dazu hat er keine Papiere - die man zur damaligen Zeit für die Reise von Göttingen nach Berlin durchaus benötigte - erzählt dem Polizeibeamten, dass Napoleon seinerzeit sogar auf eine Kanonade Göttingens nur seinetwegen verzichtet hätte...dazu muss sich der Leser die damaligen Verhältnisse in "Deutschland" vor Augen führen. Die herrschenden Fürsten hatten eine Heidenangst vor revolutionären "demokratischen" Umtrieben...und auf Napoleon war man in Deutschland nach dem Wiener Kongress in kaum einem der dazu zählenden 100+x Kleinstaaten allzu gut zu sprechen. Naja...der Polizist muss einen verdächtigen "Turner" (Jawoll....gedenke man doch auch dem "Turnvater" Jahn) verfolgen und beide Gauss gelangen schließlich wohlbehalten nach Berlin, wo sie schon von einem umtriebigen Alexander von Humboldt in Empfang genommen werden. Sehr schön auch die Schilderung, wie man mit Hilfe der noch nicht so recht ausgereiften Erfindung des Herren Daguerre versucht "die Zeit festzuhalten"....
Nun...in diesem Stil wird uns schließlich das Leben der beiden ungleichen Geistesgrößen und Sonderlingen geschildert. Während Gauss sich an die Entdeckung der Mathematik (und schließlich auch der Physik) macht und eine 'innere Welt' bis an ihre Grenzen erforscht, begibt sich Humboldt auf Entdeckungsreise nach Süd- und Mittelamerika, die 'äußere Welt' bis zu ihren Grenzen zu erforschen.
Am Ende bemerken die beiden mittlerweile schon ergrauten Herren, dass sie doch gar nicht so verschieden sind - auch wenn ihr Leben kaum unterschiedlicher hätte sein können.

Ich hab das Buch sehr genossen. Vor allem natürlich, da ich mich durch die geschilderten Eigenarten der beiden Protagonisten an den ein oder anderen hochbegabten (aber doch recht schrulligen) Zeitgenossen erinnert gefühlt habe, der meinen Weg bislang gekreuzt hat...natürlich entdeckt man auch die ein oder andere eigene 'Seltsamkeit' wieder. Ein weiteres Highlight ist für mich Kehlmanns Sprachgewalt. Nein, das Werk ist nichts für den "Wald-und-Wiesen-Gelegenheits-JerryCotton-Leser". Ganz im Gegenteil. Natürlich wirkt die Sprache (ich sage nur "lang lebe der Konjunktiv"!) etwas antiquiert, aber wir befinden uns ja schließlich in der ersten Hälfte des 19. Jahrhunderts und nicht bei RTL.
Fazit: Ich hab schon lange nicht mehr ein so kurzweiliges und interessantes Buch gelesen. Lesebefehl!

Sunday, March 04, 2007

Ning and Library 2.0


Again, it was only a matter of time, until a Web 2.0 Site turns up that allows you to create your own social networking web site. Ning hosts numerous socual networking web sites covering all possible issues. It's rather simple...just register and create your own site following the dialog.
Yesterday - again by chance...i.e. by tag-browsing bibsonomy entries - I found one of Ning's social networking sites on the topic 'library 2.0'. I signed up and I'm rather curious what will happen there :)

Friday, February 23, 2007

Internet Pioneers ... must see!


While doing some research on internet history for writing the 2nd edition of my WWW book, I found an impressive film that is gathering a lot of the most important internet pioneers. The documentary is entitled "Computer Networks: The Heralds of Resource Sharing" and was produced back in 1972 (!!!) by Steven King from MIT. It features the ARPANET and many of the most important names in the history of computer networking.
You will see J.C.R. Licklider, former director of the IPTO at ARPA, who was the first to envision a global internet. Also Larry Roberts (also former director of IPTO), Robert Kahn (co-inventor of the TCP/IP protocol and winner of the Turing Award), or Donald Davies (co-inventor of packet switching) are giving contributions.
It's really a 'must see', simply because all I already knew about internet history had come from books. I also only knew a few pictures showing early internet hardware or some mug-shots of the mentioned scientists. It's really interesting to see them in the film and to hear them talking about their great vision of internet computing as it has become reality today.....30 years after the film was produced...

Thursday, February 22, 2007

Conspiracy ahead...Dan Brown - Angels and Demons

Finally, another year after reading 'The Da Vinci Code', I decided to give its predecessor - Angels and Demons. - a try. Ok, I really liked the 'Schnitzeljagd' (scavenger hunt) of connecting seemingly unconnected facts into wild conspiracy theories as it was presented by Dan Brown it 'The da Vinci Code'...although I was disappointed by its ending. But, as most times, it's hard to put an end to a story that is trapped in an apparently ever lasting climax. Thus, I thought, maybe the predecessor would be a little bit more well balanced.
Anyway, I had high expectations.....
So...you've heard about the Illuminati? Yes, I know. Ever since Robert A. Wilson's Illuminatus Trilogy, the Illuminati have been subject to incredible conspiracy theories. This enlightenment secret society of freethinkers, most times connected to their Bavarian section founded by Adam Weishaupt back in the 18th century, where illustrious men like Goethe have been reputed members. A lot of connections have been tried to make to Freemasons or Rosicrucians, and because of their general close connection to the movement of enlightenment -- including their opposition to the church and christian faith -- as well as for being a secret society (the government always is afraid of conspiracies) they have been banned. But, there is a lot of conspiracy literature -- reputable as well as pure fiction -- where you can read all about.
I've read Wilson's Illuminatus more than 20 years ago. As being a teenager by that time, I was really fascinated that there should be a (entire different) world out there that only opens up for those who are enlighted. Everybody else was only able to see the surface and only a happy few were able to look behind the things of daily life ... although the traces were so obvious.
So, the concept of 'Angel and Demons' was not so new to me. Dan Brown tries to draw connections between modern particle physics (the plot starts with a murder taking place at CERN) and its destructive potential (antimatter and its disastrous effects), the ancient conflict between christian faith and natural science (a.k.a. the Vatican gang against the enlightened conspirers), and the moral values of faith and christianity at all. The story is driving an accelerating pace within this conflict, and a Harvard professor of semiotics together with a female (...and rather sexy) CERN scientist trying to solve the conspiracy puzzle that is threatening the Catholic church in its very foundations. For sure this scavenger hunt is rather exciting and thrilling, but also somehow frustrating.
But for me, the end (this time I won't spoil) was reconciling again (at least a little bit, although not everything was explained, as e.g. the provenance and the story of the assassin). But we learn, that there is nothing miraculous about and we don't have to be afraid of world threatening conspiracies.
Ok...I guess you have to read it by your own. It's really entertaining...and you will have a lot places to see, when you are visiting Rome and the Vatican the next time.
Before I forget, you should really read the original English version. The German translation (I've read a few pages) is rather bad. I mean, it's well translated, but the language is rather shallow. Maybe it's the same with he English version...but for being a foreign speaker, maybe I don't realize. At least it was pretty simple to read and not difficult at all.

Monday, February 19, 2007

LEARNTEC, Karlsruhe February 13-15


This year, we participated at the LEARNTEC Fair in Karlsruhe (February 13-15). LEARNTEC is focussed of e-learning technology integrating universities and industries together within an exhibition and a congress. As officially being the advisor of an ESF/BMBF funded startup company called OSOTIS, I was visiting my students who took part at this exhibition. OSOTIS is also the name of the 'Academic Video Search Engine' that serves as a testbed for our research in semantic web and multimedia search technology.

The setting of OSOTIS is the following: We are dealing with lecture recordings and offer a search service over and also inside those lecture recordings. The main advantage of OSOTIS is that most of the video post processing that is necssary for implementing a search is done in a completely automated way. Many other video search systems depend on cost intensive post processing, such as segmenting the videos into short 'learning objects', manually annotating the video segments, etc.
OSOTIS is different:
It makes use of additional information resources such as desktop presentation (e.g. powerpoint or pdf slides or simply desktop recordings) that can be synchronized with the video recording in different ways. If there is only a lecture video without any additional information source, even speech recognition technology is able to provide keywords that can be used for the video annotation. In this way, the video can be automatically segmented and the segmants can be annotated with keyword descriptors. Additionally, if there is no way to determine the content of the video, OSOTIS offers manual annotation and social tagging services to all registered users. Thus, there is always some way to search inside each lecture recording, no matter if additional information resources are available or not.
You just enter a keyword and OSOTIS will display a list of lecture recordings that are related to that keyword. By selecting one of the results, the video will start at exactly that point in time that is directly related to the user query. OSOTIS does not host the video resources on its own server, but offers only links to the original streaming servers (for streaming resources) or origin servers with podcast/videocast recordings. Thus, also all kind of video formats can be maintained, as e.g., real media, mpeg, mp4, flash video, and others.
Up to now, the main part of hosted video lectures is in given in German (and thus being hosted by german speaking universities, as also Austria or Switzerland). But, the number of lecture recordings in English will be increased soon.

Tuesday, February 06, 2007

Arthur Schnitzler - Traumnovelle (Dream Novel)


I really have to hurry up, because in reading I'm still 2 books ahead of my reviews. So, after 'Rouge et Noir' I decided to read Arthur Schnitzler's 'Traumnovelle', which b.t.w. formed also the basis for Stanley Kubrick's as 'Eyes wide shut'. I guess I read the book because of the movie, but after reading the book I must confess that I really like the book much more. So, basically it's about that couple living in Vienna. The time is about at the beginning of the 20th century. He's a physician and the novel starts when the two are about to visit a ball (same as in the movie). Home again, she tells him about her dream and an incident that happened during their last holidays. There was a stranger and she was very attracted to him...and if he (the stranger) had said only one word, she would have followed him no mater of the consequences (but of course this didn't happen). He (the physician), somehow, is really shocked by this revelation. Then, in the night, he is called into the house of a dying patient, but as he arrives, the patient has already passed away. When leaving the patient's house again, the daughter of the dead (with tears in her eyes, her fiancé waiting for her in the other room) tells him that she loves him. Very moved, he's running through the streets of Vienna, doesn't want to go home, still thinking of some kind of 'revenge' for the 'imaginary deed' of his wife. He follows a prostitute to her home but leaves her place already before coming to business. In a bar, he meets an old friend who is playing the piano. The friend tells him that he is invited to play piano at some secret (private) party, and that people there are celebrating some kind of secretive and 'orgiastic' carnival. He persuades his friend to play some trick to get him into that party, but to get in he is in need for a mask to disguise his identity. In the middle of the night, he goes to rent a mask at a shop (again another story telling about the shop keeper selling his daughter as a prostitute...). Nevertheless wearing the mask he succeeds in getting into the party, but a strange (attractive and almost nude) woman realizes his presence there and that he is not supposed to be at this place. She warns him, but he does not care. Other people realize that he is a stranger and he is asked for the password that he can't provide. He is supposed to be punished, but the strange woman takes his place and therefore also the punishment for him. The next day, when he got home, he reads about some strange killing of a noble woman taking place in a hotel that very night, and he decides to find out, whether this killing and the 'punishment' of the last night might be connected to each other.....
I won't tell you how it ends - anyway, if you have seen 'Eyes wide shut' for sure you will know. But....it was really some experience to read it and I have enjoyed it very much. Schnitzler leaves many things to your own imagination...and Kubrick for sure invested a lot of it to setting it into scene. But, as for any movie that is based on literature, the movie just shows a special reading of the book with special emphasis on things that the director regarded as being important. Kubrick did some great job....but sorry, I don't like Tom Cruise as an actor (the only movie where I really liked his performace was 'Magnolia'...but the performance of Philip Seymour Hoffman was much more impressive...). Thus, reading the book opens up new possibilities, new ways on how your imagination might put some light into the strange story. I can highly recommend it (and it's rather short..you can make it in just one day).

P.S. you might find some other works of Arthur Schnitzler at Project Gutenberg

Saturday, January 27, 2007

Stendhal - Rouge et Noir


It seems to be the time to write about the first big novel I have read this year...although I'm already 2 books ahead and otherwise I will loose track completely. As usual - and as I have read the book in German translation - I will write a short comprehension in english, but will discuss everything in German.
Stendhal a.k.a. Henri Beyle put the scenery of "Rouge et Noir" in the time of about 1830, the Bourbone restauration in France, and subtitled it as a chronicle of the 19th century - which was still young at his time. But, it was supposed to be a novel taking place right now...and not in the past. Julien Sorell, the unusual intelligent son of a simple wood cutter - at least as being a designated priest he could speak Latin and had an enormous memory that he showed when citing entire parts of the bible by heart (and in Latin) - got the job of a house teacher in the family of the local Mayor M. de Renal. He seduces Mdme. Renal - not really out of love, but more because of his ego - and to avoid a scandal he is forced to leave. He joins the priest seminar - which by the way is one of the most impressive written parts of the book - and finally succeeds in becoming the private secretary of Marquis de la Mole. The Marquis' daugther soon got an eye on Julien and finally - this really takes Julien some time and and also sophisticated strategies - they plan to marry because she became pregnat (by him...). Of course the Marquis is rather dissappointed about this misalliance. Then, he receives a letter written by Mdme. de Renal in which she warnes the Marquis de la Mole about Julien being an imposter, whose only goal is to make carreer out of seducing women in the families where he is put in. Julien also reads the letter and for revenge shoots Mdme. de Renal while she is attending at church. Although she recovers, Julien gets voluntarily adjudged and executed......

Die vorliegende neue deutsche Übersetzung von Stendhals Klassiker ""Rot und Schwarz" kann ich allen - egal ob Fan von französoscher Literatur des 19. Jahrhunderts oder nicht - nur wärmstens ans Herz legen. Das Buch ist überaus spannend und unterhaltsam geschrieben. Stendhals mitunter kurze und prägnante Art verzichtet auf ausschweifende Schilderungen der Schauplätze ohne jedoch das jeweils für diese typische außer Acht zu lassen. Üppig, intensiv und wohlüberlegt ausgefallen sind alle Dialoge. Man durchlebt die Höhen und Tiefen von Julien Sorells Dasein - auch wenn man seine Gefühle, seinen Antrieb heute nicht immer recht verstehen kann. Die französische Revolution, Napoleons Kaiserreich und die anschließende Restauration - auf die eine weitere Revolution folgen sollte - prägen das gesellschaftliche Bild, das Stendhal zeichnet. Der Karrierist und bürgerliche Emporkömmling wird ebenso scharf charakterisiert wie der alteingesessene Adel, der ewige Streit zwischen Jesuiten und Jansenisten verfolgt die Handlung wie das gerade im Entstehen begriffene Genre des Stutzers und modebewußten Dandytums. Und natürlich die Frauen...alle scheinen sie in Julien verliebt. Angefangen von der unscheinbaren Kammerzofe, über Mdme. de Renal, einer Kaffeehausangestellten, einer verwittweten Generalin, bis hin zur Marquise de la Mode...alle weiß Julien von sich einzunehmen...und zu enttäuschen.
Das Ende jedoch - laut Stendhal Bestandteil der dem Buch zugrundeliegenden wahren Begebenheit - bleibt mir rätselhaft. Wie bereits geschildert versucht Julien Mdme. de Renal in der Kirche zu ermorden und sieht danach, obwohl diese sich von ihren Verletzungen erholt und ihm vergibt, keinen anderen Ausweg, als sich dem Gericht zu überantworten und selbst auf seine Verurteilung zum Tode zu bestehen. Natürlich...nicht gerade ein 'Hollywood'-gerechtes Ende. Aber eindringlich und wirklich kurzweilig erzählt. Besonders hervorzuheben sind in dieser Ausgabe die vielen Zugaben. Neben einem ausführlichen Anhang mit Erklärungen und Anmerkungen Stendhals (die man im laufenden Text jeweils nachschlagen kann..) bietet die Ausgabe noch Entstehungs- und Wirkungsgeschichte, sowie Stendhals eigene Rezension des Werkes. Also: Lesebefehl!

Wednesday, January 24, 2007

SOFSEM - Day 4

Now we have snow....finally :) ...even a lot of it. It was snowing all day long, roads in Czech Republic and also in southern Germany were closed. Also Prague Airport was closed until the afternoon. But, I guess as far as I remember that are the more typical weather conditions for SOFSEM.
Anyway, the day started with a keynote of Tom Henziger about 'Games, Time, and Probability: Graph Models for System Design and Analysis'. He addressed three major sources of system complexity: concurrency, real time, and uncertainty. Concurrency can be modelled as a multi-player game representing a reactive system with potential collaborators and adversaries. Real time requires the system to combine discrete state changes as well as continous state evolution, while state changes - for uncertainty - also have to be modelled in a probabilistic way.
Unfortunately some of the presenters of the following contributed papers did not show up. Thus, the conference program was subject to several changes. In the afternoon the posters of the student research forum each had a short 5 minute presentation, followed by a poster exhibition and a lot of discussions. In the end, the participants should give a vote for the best poster presentation. My choice - which of course is completely subjective - was the poster of of Henning Fernau and Daniel Raible on 'Alliances in Graphs: a Complexity-Theoretic Study'.
In the late evening I was trying to look for my car, which was buried under the snow at the parking lot. Due to the wind the snow around the parking lot (and my car) was piled up almost half a meter...which made me think about the road conditions and the plan of driving home the next day....

Tuesday, January 23, 2007

SOFSEM 2007 - Day 3

Today started with a keynote given by Ricardo Baeza-Yates from Yahoo! Research on 'Mining Web Queries'. In particular he showed how to identify categories of user queries and how to use this information to create an appropriate ranking of the search results. Besides the already identified 'coarse' categories, such as, e.g., queries being 'informational', 'navigational', or 'transactional' (which means that the user wants to have (a) information about a specified topic, (b) a starting point for further research, or (c) a homepage related to the resource for transactional purposes (e.g. shopping)...), he addressed several graphs that can be compiled out of the search engine logfile, as e. g., URL cover graph, URL link graph, session graph...These graphs can be used for identifying polysemic expressions, similar or related queries, clusterings of queries, or even a (pseudo)taxonomy of queries.
Besides web query mining, he mentioned some interesting numbers concerning Yahoo, as e.g. that Yahoo administrates about 20 PetaBytes of Data with more than 10 TeraBytes of data traffic per day. But, on the other hand, he gave an estimation of the actual world knowledge and related it to the ammount of data managed by Yahoo today: given that a person creates about 10 pages of data concerning a distinct event, and if we estimate the number of events of about 5000 in a lifetime, and if we multiply that number by the world's population....we will end up with about 0,0057% of the 'world knowledge' currently being represented in Yahoo...

Monday, January 22, 2007

SOFSEM 2007 - Day 2


The second day of SOFSEM started with a keynote of Bertrand Meyer (maybe you remember Eiffel...) from ETH Zürich on 'Automatic Testing of Object-Oriented Software'. To enable automated testing, he referred the concept of 'contracts' being directly embedded in the classes of the Eiffel programming language. With a contract you are able to specify the software's expected behaviour (preconditions, postconditions, and invariants). which can be monitored during execution. In automated software testing, contracts may serve as test oracles that decide, whether a test case has passed or failed. He presented 'Auto Test' unit testing framework, which is using Eiffel contracts as test oracles. Auto Test is able to exercise all classes by generating objects and routine arguments. Also manual testing can be embedded as well as regression testing for failed test cases, which is implemented in a 'minimized' form by retaining only the relevant instructions.

For the rest of the second day contributed (refereed) paper presentations are scheduled. I will have to chair the first session of the 'emerging web technologies' track, which will be on XML technology. If there (or in any other session I attend) will be anything of interest, you will read it right here ... :)
So...Joe Tekli from the Université de Bourgogne presented a 'Hybrid Approach on XML-Similarity', which combined structural similarity of XML-Documents with 'semantic' arguments, i.e. tag names of different XML-documents are compared with the help of WordNet to compute some similarity measure. Quite an interesting application that can be build on, esp. regarding the semantic similarity aspect. But nevertheless, maybe we can use it for our MPEG-7 based video search system (OSOTIS).

Sunday, January 21, 2007

SOFSEM 2007 - Day 1


This year, after about 7 or 8 years, I am attending again the SOFSEM conference on 'Current Trends in Theory and Practice of Computer Science' (for the 2nd time). Maybe SOFSEM is not the most important of all the computer science conferences around, but it is rather original and has quite some history (i.e. it's tradition dates back more than 30 years...). SOFSEM means SOFtware SEMinar, and this already gives some hint about its originality. Starting from a winter lecture with only limited international attendance it has developed to an interesting mixture of lectures (given by invited speakers of significant reputation), presentations of reviewed research papers, and student paper presentations. By tradition, it's location always switches between somewhere in Slowakia and the Czech Republique and always in winter. Unfortunately, this year winter did not really show up and thus, we are sitting here in Harrachow (a well known winter resort) without any snow. On the other hand, nice thing about this situation is that travelling this year has become much easier (because there is no snow even in the mountain areas).
This year, I am co-chairing the track 'emerging web technologies' as being one of the four SOFSEM tracks. By tradition, there is always a track 'foundations of computer science' besides of three changable tracks concering breaking topics of current interest , i.e. (in this year) 'multi-agent systems', 'emerging web technologies', and 'dependable software and systems'.
The first day on SOFSEM, after the opening note given by Jan van Leeuwen, in which he referred to the long tradition of SOFSEM and to Czech computer science history, starts with a full day of invited lectures covering all four topics.

  • Manfred Broy from TU Munich started with a presentation on 'Interaction and Realizability'. In interactive computation - in difference to sequential, atomic computation - input as well as output is not provided as a whole, but step by step while the computation continues. He pointed out that interactive behaviour can be modeled with Moore machines and introduced the term of 'realizability', which is a fundamental issue when asking whether a behaviour corresponds to a computation. 'Realizable functions' are defined as being abstractions of state machines (in a similar way as partial functions are abstractions of Turing machines) and can be used to extend the idea of computability to interactive computations.

  • Andrew Goldberg followed with a talk on 'Point-to-Point Shortest Path Algorithms with Preprocessing'. To run on even small devices while at the same time covering graphs with tens of millions of nodes (as, e.g., in roadmaps for navigation devices), off course efficient algorithms are required. The traditional way is to search a ball around the starting point (as e.g. in Dijkstra's algorithm) that can be speed up by biasing the search towards to intendet target point (as e.g. in A* search, if additional information is available that provides a lower-bound on the distance to the target) or by pruning the search graph (as e.g. in ALT algorithms that precompute distances to preselected landmarks, or using 'reaches').

  • Jerome Lang from IRIT (France) continued the afternoon session with a survey on 'Computational Issues in Group Decision Making', which combines 'social choice' (from economics) and AI (applications) into 'computational social choice' theory. In this new and very active discipline concepts as e.g. voting procedures, coalition formation, and fair division (from social choice), which is also important for multi-agent systems, are examined under the consideration of complexity analyses and algorithm design.

  • I realized that I will be the chairman of today's last session. Thus, the summary of Remco Veltkamp's (University of Utrecht, The Netherlands) talk on 'Multimedia Retrieval Algorithms' will come with a little delay....
    The presentation started with citing Marshal McLuhan's famous quote 'The medium is the message' smartly being connected to the basic definitions of multimedia retrieval. Difficult thing in multimedia retrieval is the proper understanding of the mechanisms of human perception and in connection to that the question of how to take care of it's peculiarity in information retrieval. E. g., the human visual system is famous for 'generic interpretations', i.e. sometimes we see things that are not really there, as already has been described by Wertheimer's Gestalttheorie back in 1923. Interesting fact, that some of these visual illusions do also exist for audio perception. For multimedia retrieval metrics have to be defined for computing similarities (as well as differences of multimedia objects) in an efficient way, while the algorithms dealing with multimedia retrieval have to be carefully designed according to the type of problem that is addressed (e.g., computing problem, optimization problem, decision problem, etc.). The presentation closed with a short demonstration of the music search engine Muugle that realizes the concept of 'query-by-humming'.

Saturday, January 13, 2007

...against all odds


On wednesday I attended a talk given by Michael Strube from EML Research on "World Knowledge induced from Wikipedia - A New Prospect of Knowledge-Based NLP ". He was showing how the (meanwhile famous) collaborative encyclopedia can be used for information retrieval purposes in a way similar to (more traditional) online dictionaries as e.g. WordNet and - though being not well structured - provides results of almost equal quality.
First thing was that for their work, Strube and his colleague regarded each Wikipedia page as being the representation of a concept (we already had some arguments about that as you might remember...). Next, they developed some metric for similarity of concepts w.r.t. to the concept hierarchy (where the wikipedia defined 'concepts' come into play). Since 2004, wikipedia features a user defined concept hierarchy. This hierarchy of concepts also can be regarded as being a folksonomy, simply because this is not a knowledge representation carefully designed by some designated domain expert, but by the wikipedia comunity in a collaborative way. Unfortunately, the wikipedia concept hierarchy suffers exactly from that fact. From my pont of view it seems problematic to compare the proposed similarity measure (based on wikipedia concept hierarchy) with other similarity measures (based on commonly shared expert ontologies). O.k., you might argue that indeed the wikipedia concept hierarchy IS commonly shared, because it has been developed by the wikipedia community...but is the knowledge represented in wikipedia really 'common'? Just remember the diversity and manifold of Star Wars characters or Star Trek episodes in wikipedia compared with, as e.g., the history of smaller Eropean countries. As for all ontologies always the view and the knowledge of the ontology designer has to be considered. The wikipedia concept hierarchy - although partly being really appropriate - reminds me somehow to this famous literary chinese dictonary entry defining the term 'animal' which is quoted by Jorge Luis Borges. Another problem lies in the fact that the different language versions of wikipedia have developed different concept hierarchies (sic!).

In the end, I was asking how this proposed information retrieval based on wikipedia could be improved by considering a 'Semantic Wikipedia', as e.g., the Semantic MediaWiki (given that those semantic wikipedias would contain sufficient data). Instead of answering my question, Michael Strube cited Peter Norwig's argument against the Semantic Web from last years AAI2006. Just to sum up: the semantic web will not become reality because of the inability of its users to provide correct semantic annotations. But hey...this guy (Strube) was talking about wikipedia. Doesn't this argument raise any associations? Just remember the time 5 or 10 years ago. Nobody (well almost nobody) would have believed that it will be possible to write an entire encyclopedia collaboratively on an open source basis - just because the web user's did not seem to be able to write 'correct' articles....

Sunday, January 07, 2007

Sam Bourne - Die Gerechten....


Etwas verspätet, aber bevor ich schon wieder das nächste Buch beendet habe, muss ich heute doch noch ein paar Worte über meine 'Feiertagslektüre', "Die Gerechten" von Sam Bourne verlieren:
Meine Erwartungen waren ja recht hoch gesteckt. Neugierig gemacht durch einen Beitrag der Kulturzeit wollte ich diesen als ungewöhnlichen Thriller angekündigten Roman unbedingt lesen. Will Monroe, ein junger Journalist der New York Times, deckt eine Spur ungewöhnlicher Morde auf, die erst auf dem zweiten Blick tatsächlich miteinander zusammenhängen. Er wird - zumindest bleibt er (und der Leser) zunächst in diesem Glauben - mit hineingezogen in eine jüdische-konservative (sic!) Weltverschwörung, deren Ziel in der Herbeiführung des Endes der Welt und im (damit erzwungenen) Erscheinen des langerwarteten Messias zu bestehen scheint. [ACHTUNG: SPOILER-WARNUNG] Fast bis zum Schluss wird man in diesem Glauben gelassen, aber 'natürlich' waren es dann doch irgendwelche sektiererischen, 'bibeltreuen' und erzkonservative Christen, die übrigens das gleiche Ziel verfolgen.[ENDE: SPOILER-WARNUNG]
Ehrlich gesagt, ich war enttäuscht. Nicht nur, dass der ganze Plott mehr als an den Haaren herbeigezogen scheint. Nein, eigentlich eher die mangelnde erzählerischen Qualitäten des Autors haben mich etwas verärgert. Die Personen werden stereotypisch (langweilig) und absolut vorhersehbar charakterisiert. Jeglicher Tiefgang - sei es der Vater-Sohn Konflikt der Hauptfigur, seine Beziehungsprobleme oder der Umgang mit seiner Ex-Freundin - erscheinen irgendwie vollkommen platt. Eigene Gedankengänge, die einem die Beweggründe der handelnden Personen näherbringen würden, fehlen fast völlig. Sicher, die Story enthält zahlreiche Cliffhanger und wird daher für eine Vielzahl der Leser spannende Lektüre bieten, aber der Handlungsfluss ist fast vollständig linear, einige Fragen bleiben ungeklärt, und die pseudo-wissenschaftlichen Herleitungen [ich sage nur: Kabbala und GPS...] verärgern den Leser eher als dass diese ihm ein Aha-Erlebnis bieten würden.

Fazit: Eine neue (alte) Verschwörungsgeschichte, die mit Endzeit-Phantasien, oberflächlich religiösen Vorurteilen und Kabbala-Techno-Babbel versucht, Boden gut zu machen, deren erzählerische Qualität meines Erachtens nach aber zu Wünschen übrig lässt...

P.S. Jedes neue Buch zum Thema Verschwörungstheorien muss sich gegen die zwei Meilensteine dieses Genres messen: Sear und Wilson's Illuminatus Trilogie (wenn schon, dann wenigstens absolut abgedreht....) oder man nimmt sich gleich den Großmeister Umberto Eco vor mit seinem Foucaultschen Pendel, in dem jegliche Verschwörungstheorien auf äußerst intelligente und lesenswerte Weise durchexerziert und ad absurdum geführt werden.

Saturday, December 30, 2006

Neal Stephenson - The Confusion (Vol. 2 of the Baroque Cycle)


I have finished 'Confusion' ... more than a thousand pages and what a story :)
As I have read the book in German, the short review will also be in German....

Ok, ich bin also durch. Erst einmal ganz dickes Lob an die beiden Übersetzer Nikolaus Stingl und Juliane Gräbener-Müller. Die schwer wiegenden 1000 Seiten kommen mit einer Leichtigkeit daher, die dieses Buch zu einem wahren Lesevergnügen werden lassen. Das Buch - eigentlich sind es ja streng genommen nach der Vorrede des Autors zwei - wartet mit zahlreichen Nebenhandlungen, erzählerischen Beschreibungen und Schnörkeln und abstrusen Handlungsverläufen auf, dass man es wahrhaft als 'barock' bezeichnen mag. Verquickt (confused) werden dabei zwei Haupthandlungsstränge - die Verschwörer (eigentlich Galeerensklaven) rund um Jack Shaftoe und ihrer Jagd nach Salomos Gold (das sie sich schnell wieder abjagen lassen) und ihrer merkantilen Expedition (-> Quecksilber) rund um den Globus....sowie die Geschichte um Eliza (mittlerweile Herzogin von Arcachon (sic!) und Qwghlm) und dem (historischen) Entstehen des bargeldlosen Zahlungsverkehrs und dessen Auswirkungen auf Wirtschaft und Politik (Confusion!) des damaligen Europas sowie Eliza's Rache an ihren ehemaligen (Duc d'Arcachon sr.) und aktuellen (von Hacklheber) Peinigern.
Etwas kurz geraten (im Gegensatz zum ersten Band 'Quicksilver') ist die Geschichte um Daniel Waterhouse (einschließlich Newton, Leibnitz, etc...), die Royal Society und den wissenschaftlichen Erkenntnissen der damaligen Zeit. Dafür stehen in diesem Band die wirtschaftlichen Errungenschaften des ausgehenden 17. und beginnenden 18. Jahrhunderts im Mittelpunkt, wie z.B. das globale Agieren der Holländischen Ostindischen Handelsgesellschaft, Spaniens Plünderung der Neuen Welt oder das Entstehen des bargeldlosen Handels.

Was mir an Stephensons Romanen immer wieder gefällt sind seine Schilderungen vollkommen abstruser Situationen, in die mitunter wichtige Dialoge und Handlungen eingebettet werden, sei es die großmaßstäbliche (behelfsmäßige) Herstellung von Phosphor aus Urin im indischen Hinterland oder die Jagd nach einer Fledermaus im Esszimmer mit Hilfe des angerosteten Rapiers von Leibnitz.
Auf alle Fälle freue ich mich schon auf den letzten Teil der Trilogie, auf dessen Übersetzung wir wahrscheinlich noch ein gutes Jahr warten werden müssen...

Links: -> Clearing up the Confusion in wired news