Showing posts with label semantic web. Show all posts
Showing posts with label semantic web. Show all posts

Thursday, December 04, 2014

The Fact Ranking Challenge continues....

Yes, you might remember our experiment with fact ranking [1,2,3,4]. After almost 4 months, this is the current state of intermediate results (cf. below).

To make it short: YES, there has been some progress. NO, it is not sufficient so far.

Therefore, PLEASE HELP!
If you have already started to work with the application [3], PLEASE CONTINUE!
If you don't know the application [3], PLEASE START!
PLEASE PARTICIPATE!



We know that is is not an easy task and we also know that it takes time. Therefore, WE REALLY DO APPRECIATE YOUR HELP VERY MUCH!!

Please keep on playing and help us to gather more data!
Please tell all your family and friends to support us!
Please tell all your colleagues and fellow students to support us!

THANK YOU VERY MUCH!

[1] The Fact Ranking Quiz Application, http://s16a.org/fr/
[2] Help us with a Research Problem, July 30, 2014
[3] The Importance of Relevance - Intermediate Results, Aug 19, 2014
[4] More intermediate Results from our Fact Ranking Challenge, Sep. 2, 2014


...and here are the statistics:

Number of users who participated: 465
Sum of concepts done: 2410
485 unique concepts are covered (out of 541). 
Average concepts done per user: 5.183

295 times concepts were skipped (relevance in Step2 hasn't been changed for any of the facts).

CONCEPTS DONE:
0 concepts were done by 97 users. 
1 concepts were done by 93 users. 
2 concepts were done by 78 users. 
3 concepts were done by 49 users. 
4 concepts were done by 33 users. 
5 concepts were done by 26 users. 
6 concepts were done by 10 users. 
7 concepts were done by 7 users. 
8 concepts were done by 9 users. 
9 concepts were done by 8 users. 
10 concepts were done by 8 users. 
11 concepts were done by 6 users. 
12 concepts were done by 2 users. 
13 concepts were done by 1 users. 
14 concepts were done by 5 users. 
15 concepts were done by 3 users. 
16 concepts were done by 3 users. 
17 concepts were done by 1 users. 
18 concepts were done by 1 users. 
19 concepts were done by 2 users. 
20 concepts were done by 3 users. 
21 concepts were done by 2 users. 
23 concepts were done by 1 users. 
24 concepts were done by 2 users. 
25 concepts were done by 1 users. 
26 concepts were done by 1 users. 
31 concepts were done by 1 users. 
36 concepts were done by 1 users. 
40 concepts were done by 1 users. 
42 concepts were done by 2 users. 
56 concepts were done by 1 users. 
58 concepts were done by 1 users. 
60 concepts were done by 1 users. 
64 concepts were done by 1 users. 
68 concepts were done by 1 users. 
70 concepts were done by 1 users. 
88 concepts were done by 1 users. 
201 concepts were done by 1 users. 

EDUCATION:
highschool : 47 users.
phd : 71 users.
other : 26 users.
bachelors : 105 users.
masters : 213 users.

AGE:
33+ :  261 users.
19-25 :  72 users.
26-32 :  122 users.
0-18 :  7 users.

GENDER:
female :  111 users.
male :  351 users.

COUNTRY OF ORIGIN:
Angola :  1 users.
Belarus :  1 users.
Portugal :  4 users.
Philippines :  2 users.
Morocco :  3 users.
Greece :  5 users.
Ukraine :  3 users.
Indonesia :  4 users.
Afghanistan :  2 users.
Sri Lanka :  1 users.
Italy :  15 users.
Iraq :  2 users.
India :  44 users.
France :  14 users.
Denmark :  1 users.
Latvia :  1 users.
Pakistan :  4 users.
Syrian Arab Republic :  2 users.
Montenegro :  1 users.
Armenia :  1 users.
Mexico :  2 users.
Canada :  1 users.
Brazil :  10 users.
Venezuela :  1 users.
Croatia :  2 users.
Macedonia, The Former Yugoslav Republic of :  1 users.
Romania :  2 users.
Western Sahara :  1 users.
Algeria :  5 users.
Sweden :  2 users.
United States :  24 users.
Serbia :  10 users.
Nigeria :  3 users.
Estonia :  1 users.
Spain :  9 users.
Taiwan, Republic of China :  2 users.
Ireland :  1 users.
Russian Federation :  12 users.
Israel :  1 users.
Colombia :  3 users.
Switzerland :  2 users.
Azerbaijan :  2 users.
Kenya :  2 users.
Yemen :  1 users.
Malaysia :  2 users.
Viet Nam :  1 users.
Australia :  5 users.
Peru :  1 users.
Albania :  1 users.
South Africa :  2 users.
Tunisia :  1 users.
Netherlands :  9 users.
China :  3 users.
Somalia :  1 users.
Slovenia :  1 users.
Finland :  3 users.
Lithuania :  1 users.
Austria :  7 users.
Sudan :  1 users.
United Kingdom :  15 users.
Egypt :  2 users.
Bahamas :  1 users.
Hungary :  1 users.
Belgium :  1 users.
Poland :  6 users.
Iran, Islamic Republic of :  2 users.
Bulgaria :  3 users.
Norway :  1 users.
Germany :  177 users.
New Zealand :  3 users.

COUNTRY OF RESIDENCE:
United Arab Emirates :  1 users.
null :  0 users.
Belarus :  1 users.
Portugal :  2 users.
Philippines :  1 users.
Morocco :  2 users.
Greece :  3 users.
Ukraine :  1 users.
Indonesia :  3 users.
Luxembourg :  1 users.
Sri Lanka :  1 users.
Italy :  14 users.
Iraq :  1 users.
India :  32 users.
France :  17 users.
Jordan :  1 users.
Denmark :  1 users.
Latvia :  1 users.
Pakistan :  3 users.
Syrian Arab Republic :  1 users.
Oman :  1 users.
Turkey :  1 users.
Czech Republic :  1 users.
Armenia :  1 users.
Canada :  4 users.
Brazil :  9 users.
Croatia :  1 users.
Romania :  3 users.
Algeria :  6 users.
Sweden :  3 users.
United States :  25 users.
Serbia :  5 users.
Nigeria :  1 users.
Saudi Arabia :  1 users.
Estonia :  2 users.
Spain :  7 users.
Taiwan, Republic of China :  1 users.
Ireland :  3 users.
Russian Federation :  6 users.
Israel :  2 users.
Colombia :  2 users.
Switzerland :  8 users.
Azerbaijan :  2 users.
Kenya :  2 users.
Norfolk Island :  1 users.
Yemen :  1 users.
Malaysia :  2 users.
Australia :  7 users.
Peru :  1 users.
Albania :  1 users.
South Africa :  2 users.
Netherlands :  13 users.
China :  1 users.
Somalia :  1 users.
Slovenia :  1 users.
Gambia :  1 users.
Finland :  3 users.
Lithuania :  1 users.
Austria :  6 users.
United Kingdom :  17 users.
Egypt :  1 users.
Bahamas :  1 users.
Belgium :  3 users.
Poland :  7 users.
Singapore :  2 users.
Iran, Islamic Republic of :  1 users.
Bulgaria :  3 users.
Norway :  1 users.
Germany :  197 users.
New Zealand :  3 users.

Overall average confidence of users about the seen concepts: 2.616

NO. OF USERS PER CONCEPT:
1 users for 67 concepts
2 users for 123 concepts
3 users for 108 concepts
4 users for 86 concepts
5 users for 73 concepts
6 users for 37 concepts
7 users for 11 concepts
8 users for 7 concepts
9 users for 2 concepts
10 users for 2 concepts
18 users for 1 concepts
On AVERAGE there are 3.398 users per concept.

NO. OF ANSWERS PER CONCEPT (STEP 1):
On AVERAGE there are 5.002 answers per concept.

NONSENSE STATEMENTS:
TOTAL number of nonsense sentences = 1771



Tuesday, October 14, 2014

And now for something completely different...

As always in October, lectures are starting again. Like every year, I will give a lecture on Semantic Web Technologies. BTW, I have realized that I give now courses on Semantic Web for almost 10 years. It all started as a seminar at the Friedrich-Schiller Universität back in Jena and became a fully-fledged lecture here at the HPI in Potsdam. Like the lecture of last winter semester, almost all lectures have been recorded and are online available either at tele-Task or yovisto.

Moreover, we have also prepared two MOOC courses Semantic Web Technologies in Spring 2013, and Knowledge Engineering with Semantic Web Technologies in Spring 2014, both very successful with thousand(s) of students.

This semester, I have decided not to do the very same all over again and to try out something completely different...

Have you ever heard of the Flipped Classroom concept? This semester, we are going to turn the lecture situation around for the students. All the lecture content has already been recorded. Thus, students can prepare for each lecture at home by watching the videos and studying the handouts as well as the course materials. Then, in the classroom, I will not present the content again, but we are going to discuss

  • everything which needs more attention according to the students,
  • everything that the students did not quite well understand,
  • including all problems, errors, and complements that seem to be important.
 Thus, to follow the (live) lecture the students have to prepare accordingly. Of course this will only work with the active participation of the students. On the other hand, it will also be more challenging for the lecturer and the tutors, because we have to be very well prepared to deal with all kind of potential questions and problems. Of course we will work out problem solutions and answers always together with the students. And it will be also the students who will take over the lead - well of course under the lecturer's guidance.

I'm very curious whether this concept will work out well with my lecture here at HPI. Please keep your fingers crossed and I will keep you posted.

Additional Links:

Tuesday, September 02, 2014

More intermediate Results from our Fact Ranking Challenge

Fact Ranking Challenge
Again, here is an update about our currently gathered data about our fact ranking experiment.

We started the original challenge about 5 weeks ago and now are able to present you some more intermediate results [1]. Nevertheless, the challenge is still running. Therefore, please distribute, participate, advertise, and help us to generate a fully fledged ground truth for fact ranking [2]. All data will be made publicly available for further research.

To determine the importance of a fact is of utmost importance, if you want to properly understand the content of information. Usually, you have a rich variety of possible interpretations of information. To determine the proper interpretation, you are going to use the context, i.e. further available information. So, the question develops from "what is important?" to "what is important with regard of this specific context?".

Current Intermediate Statistics (from Aug 28):

Number of users who participated: 388

Sum of concepts done: 1736

446 unique concepts are covered (out of 541). 

Average concepts done per user: 4.47

CONCEPTS DONE:
0 concepts were done by 79 users. 
1 concepts were done by 79 users. 
2 concepts were done by 68 users. 
3 concepts were done by 41 users. 
4 concepts were done by 25 users. 
5 concepts were done by 22 users. 
6 concepts were done by 8 users. 
7 concepts were done by 7 users. 
8 concepts were done by 9 users. 
9 concepts were done by 6 users. 
10 concepts were done by 8 users. 
11 concepts were done by 4 users. 
12 concepts were done by 2 users. 
13 concepts were done by 1 users. 
14 concepts were done by 4 users. 
15 concepts were done by 3 users. 
16 concepts were done by 2 users. 
19 concepts were done by 1 users. 
20 concepts were done by 3 users. 
21 concepts were done by 2 users. 
22 concepts were done by 1 users. 
23 concepts were done by 1 users. 
24 concepts were done by 2 users. 
25 concepts were done by 1 users. 
31 concepts were done by 1 users. 
38 concepts were done by 2 users. 
40 concepts were done by 1 users. 
42 concepts were done by 1 users. 
58 concepts were done by 1 users. 
59 concepts were done by 1 users. 
62 concepts were done by 1 users. 
64 concepts were done by 1 users. 

EDUCATION:
highschool : 41 users.
phd : 57 users.
other : 22 users.
bachelors : 93 users.
masters : 174 users.

AGE:
33+ :  216 users.
19-25 :  68 users.
26-32 :  98 users.
0-18 :  5 users.

GENDER:
female :  90 users.
male :  297 users.

COUNTRY OF ORIGIN:
Angola :  1 users.
Belarus :  1 users.
Portugal :  4 users.
Philippines :  2 users.
Morocco :  3 users.
Greece :  5 users.
Ukraine :  3 users.
Indonesia :  3 users.
Sri Lanka :  1 users.
Italy :  13 users.
Iraq :  1 users.
India :  40 users.
France :  11 users.
Latvia :  1 users.
Pakistan :  3 users.
Syrian Arab Republic :  1 users.
Montenegro :  1 users.
Armenia :  1 users.
Mexico :  2 users.
Brazil :  10 users.
Venezuela :  1 users.
Croatia :  2 users.
Macedonia, The Former Yugoslav Republic of :  1 users.
Romania :  2 users.
Western Sahara :  1 users.
Algeria :  4 users.
Sweden :  2 users.
United States :  21 users.
Serbia :  8 users.
Nigeria :  2 users.
Estonia :  1 users.
Spain :  8 users.
Taiwan, Republic of China :  2 users.
Ireland :  1 users.
Israel :  1 users.
Russian Federation :  9 users.
Colombia :  3 users.
Switzerland :  1 users.
Azerbaijan :  1 users.
Kenya :  2 users.
Yemen :  1 users.
Malaysia :  2 users.
Viet Nam :  1 users.
Australia :  4 users.
Peru :  1 users.
Albania :  1 users.
South Africa :  2 users.
Netherlands :  8 users.
China :  2 users.
Somalia :  1 users.
Slovenia :  1 users.
Finland :  3 users.
Lithuania :  1 users.
Austria :  6 users.
Sudan :  1 users.
United Kingdom :  15 users.
Egypt :  2 users.
Bahamas :  1 users.
Hungary :  1 users.
Poland :  4 users.
Iran, Islamic Republic of :  2 users.
Bulgaria :  3 users.
Norway :  1 users.
Germany :  140 users.
New Zealand :  3 users.

COUNTRY OF RESIDENCE:
United Arab Emirates :  1 users.
Belarus :  1 users.
Portugal :  2 users.
Philippines :  1 users.
Morocco :  2 users.
Greece :  3 users.
Ukraine :  1 users.
Indonesia :  2 users.
Luxembourg :  1 users.
Sri Lanka :  1 users.
Italy :  11 users.
India :  31 users.
France :  13 users.
Jordan :  1 users.
Denmark :  1 users.
Latvia :  1 users.
Pakistan :  3 users.
Oman :  1 users.
Turkey :  1 users.
Czech Republic :  1 users.
Armenia :  1 users.
Canada :  3 users.
Brazil :  9 users.
Croatia :  1 users.
Romania :  3 users.
Algeria :  4 users.
Sweden :  3 users.
United States :  23 users.
Serbia :  4 users.
Nigeria :  1 users.
Saudi Arabia :  1 users.
Estonia :  2 users.
Spain :  7 users.
Taiwan, Republic of China :  1 users.
Ireland :  3 users.
Israel :  2 users.
Russian Federation :  2 users.
Colombia :  2 users.
Switzerland :  7 users.
Azerbaijan :  2 users.
Kenya :  2 users.
Norfolk Island :  1 users.
Yemen :  1 users.
Malaysia :  2 users.
Australia :  6 users.
Peru :  1 users.
Albania :  1 users.
South Africa :  1 users.
Netherlands :  12 users.
Somalia :  1 users.
Slovenia :  1 users.
Gambia :  1 users.
Finland :  3 users.
Lithuania :  1 users.
Austria :  5 users.
United Kingdom :  17 users.
Egypt :  1 users.
Bahamas :  1 users.
Belgium :  2 users.
Poland :  5 users.
Singapore :  1 users.
Iran, Islamic Republic of :  1 users.
Bulgaria :  3 users.
Norway :  1 users.
Germany :  154 users.
New Zealand :  2 users.

Overall confidence of users about the seen concepts: 2.687

NO. OF USERS PER CONCEPT:
On AVERAGE there are 2.18 users per concept.

NO. OF ANSWERS PER CONCEPT (STEP 1):
On AVERAGE there are 4.71 answers per concept.

NONSENSE STATEMENTS:

TOTAL number of nonsense sentences = 1371


Hint: You might wonder about the impressive high scores on the top of the list? Well, actually points are given exponentially, i.e. the longer you play, the more points you will score per processed concept.

References:
[1] Help Us with a Research Problem, July 30, 2914
[2] Fact Ranking Web-Application, http://s16a.org/fr/

Tuesday, August 19, 2014

The Importance of Relevance - Intermediate Results

Current Highscore List of our fact ranking challenge
In my last post, we invited you to take part in our research challenge, which was about creating a ground truth for fact ranking algorithms. To determine the importance of a fact is of utmost importance, if you want to properly understand the content of information. Usually, you have a rich variety of possible interpretations of information. To determine the proper interpretation, you are going to use the context, i.e. further available information. So, the question develops from "what is important?" to "what is important with regard of this specific context?".

We started the original challenge about 3 weeks ago and now are able to present you first intermediate results [1]. Nevertheless, the challenge is still running. Therefore, please distribute, participate, advertise, and help us to generate a fully fledged ground truth for fact ranking [2]. All data will be made publicly available for further research.

Current Intermediate Statistics:
Number of users who participated: 110 (Thanks to all you you!!!)
Number of overall processed concepts: 509
Overall 200 unique concepts are covered (out of 541).

Average concepts processed per user: 4.63

Detailed number of processed concepts per user:
0 concepts were done by 15 users.
1 concepts were done by 27 users.
2 concepts were done by 21 users.
3 concepts were done by 12 users.
4 concepts were done by 7 users.
5 concepts were done by 8 users.
6 concepts were done by 2 users.
7 concepts were done by 1 users.
8 concepts were done by 4 users.
9 concepts were done by 1 users.
10 concepts were done by 2 users.
11 concepts were done by 2 users.
14 concepts were done by 1 users.
16 concepts were done by 1 users.
20 concepts were done by 1 users.
22 concepts were done by 1 users.
25 concepts were done by 1 users.
31 concepts were done by 1 users.
53 concepts were done by 2 users.

Participant statistics:

EDUCATION:
highschool : 7 users.
bachelors : 28 users.
masters : 47 users.
phd : 23 users.
other : 5 users. 

AGE:
33+ : 43 users.
26-32 : 34 users.
19-25 : 33 users.

GENDER:
female : 25 users.
male : 85 users.

COUNTRY OF ORIGIN:
United States : 9 users.
Serbia : 6 users.
Spain : 2 users.
Ukraine : 1 users.
Russian Federation : 3 users.
Colombia : 1 users.
Italy : 6 users. India : 2 users.
France : 3 users.
Malaysia : 1 users.
Australia : 1 users.
Albania : 1 users.
China : 1 users.
Pakistan : 1 users.
Finland : 1 users.
Austria : 2 users.
Montenegro : 1 users.
United Kingdom : 8 users.
Brazil : 3 users.
Poland : 2 users.
Iran, Islamic Republic of : 1 users.
Macedonia, The Former Yugoslav Republic of : 1 users.
Croatia : 1 users.
Germany : 49 users.
Algeria : 1 users.
New Zealand : 1 users.
Sweden : 1 users.

Overall confidence of users about the seen concepts: 2.585

NO. OF USERS PER CONCEPT:
On AVERAGE there are 1.295 users per concept.

NO. OF ANSWERS PER CONCEPT (STEP 1):
On AVERAGE there are 4.825 answers per concept.

We will keep you posted about the results.
Please distribute, participate, advertise, and help us to generate a fully fledged ground truth for fact ranking.

Hint: You might wonder about the impressive high scores on the top of the list? Well, actually points are given exponentially, i.e. the longer you play, the more points you will score per processed concept.

References:
[1] Help Us with a Research Problem, July 30, 2914
[2] Fact Ranking Web-Application, http://s16a.org/fr/

Wednesday, July 30, 2014

Help Us with a Research Problem

As you might know, we already have tried previously to let the public participate in our research. Last time, we have had developed several games (with a purpose). This time, unfortunately it is not a game, simply because the development of a good game is really expensive. But, let's get to the point. What is the task all about, where you can help us....?

You know, my research group is working on semantic technologies. Semantics in that sense means, we are trying to (automatically) understand what information (or data) is all about and what is the meaning of it. Sometimes, information is ambiguous. This makes it difficult to understand, because you have to solve ambiguities with the help of context.

On the other hand, sometimes you have various different information about a subject. How do you determine, which information or fact is more important or relevant than another? Just a quick example. Let's assume we have the following two facts:

(1) Albert Einstein is a physicist.
(2) Albert Einstein is a Vegetarian.

Which of the two facts is more important or relevant? Yes, this is difficult to answer, simply because the truth often lies in the eye of the beholder. For a vegetarian, maybe the second fact is more important. But, what about the most common opinion? What would the mainstream think? Probably, most people would say that fact (1) in general is more important.

So, what we are doing is that we develop heuristics that determine the importance of facts (relative to other facts). To get an idea about the quality of our heuristics, we have to do an evaluation, i.e. somebody has to decide whether the decision of the heuristics was wrong or right. Unfortunately, there does not exist a ground truth for this task called "fact ranking". Therefore, we are about to create a new ground truth that later will be publicly available and open for all researchers.

This ground truth is achieved with the little 'voting' application that you will find here [1]. You just have to register with the tool and then the task will be explained to you in detail. We took 500 popular concepts from Wikipedia and you have (1) to think about the most important facts about these concepts that come to your mind and then (2) rate the (new) facts presented to you according to their relevance. There is no right or wrong answer. Just vote as you think it seems right for you. Afterwards, we will aggregate all votes from all participants to determine the general (mainstream) relevance of the presented facts.

You might interrupt your rating of the presented facts at any time you like and continue later. To make it a bit more interesting, you can also score points and of course there is a highscore list. We would really appreciate your help in this task. Please do also spread the word. The more participants, the more valid our ground truth will be.

We know that this is a difficult and sometimes rather boring task. The more we would be really grateful for your assistance!

[1] Fact Ranking Web-Application, http://s16a.org/fr/

Friday, June 20, 2014

Harald's Original Miscellany - Prolificacy vs. Popularity in Literature

Oh wow, it's quite a while that I wrote my last post here in the blog... But, while preparing exercises for the OpenHPI MOOC course 'Knowledge Engineering with Semantic Technologies', I was about to play around a little with SPARQL to come up with new exercises for the students of the course. To make it short, our current lecture examples all deal with writers and books. Thus, to learn how to query RDF knowledge bases with the SPARQL query language, I chose the DBpedia. While trying to think of some interesting toy examples, I started to play around and the facts that I discovered by chance were so interesting that I totally forgot about my lunch break :)

So here are some interesting facts about books and authors that will be continued in later posts. All presented statistics is based on the (English) Wikipedia (of course for the SPARQL queries we use DBpedia)...but nevertheless, it is wikipedia knowledge.

There are currently 15,328 authors listed (i.e. they are member of the class dbpedia-owl:Writer). First thing I wanted to find out was, who are the most prolific authors according to Wikipedia (at least this means, whose works also exist as Wikipedia Pages and who are connected via dbpedia-owl:author).

Well, here are the Top 40 Most Prolific Writers:

name numOfWorks popularityOfWorks
"L. Sprague de Camp"@en 128 10.9
"Agatha Christie"@en 103 32.7
"Isaac Asimov"@en 75 28.3
"Stephen King"@en 75 44.0
"Philip K. Dick"@en 74 14.4
"Edgar Rice Burroughs"@en 73 18.7
"Ruth Rendell"@en 70 5.3
"Dean Koontz"@en 67 6.2
"Lin Carter"@en 64 9.0
"Terry Pratchett"@en 63 41.7
"Jules Verne"@en 63 31.2
"P. G. Wodehouse"@en 63 20.9
"Robert E. Howard"@en 62 10.5
"Gary Paulsen"@en 61 4.1
"K. A. Applegate"@en 61 18.5
"August Derleth"@en 60 5.2
"John Dickson Carr"@en 59 5.9
"H. G. Wells"@en 57 30.0
"James Patterson"@en 56 12.0
"Leslie Charteris"@en 55 8.8
"Robert A. Heinlein"@en 52 36.5
"Rex Stout"@en 52 21.4
"Arthur C. Clarke"@en 50 20.8
"Harry Turtledove"@en 49 11.3
"Ray Bradbury"@en 49 15.3
"David Weber"@en 49 30.9
"Danielle Steel"@en 48 3.3
"J. M. G. Le Clézio"@en 48 3.7
"Henry James"@en 48 18.5
"Piers Anthony"@en 48 14.0
"Clive Cussler"@en 47 10.1
"Roger Zelazny"@en 45 10.5
"Alan Dean Foster"@en 44 6.7
"Joe R. Lansdale"@en 43 4.7
"Gordon R. Dickson"@en 41 5.0
"Marion Zimmer Bradley"@en 41 7.6
"Samuel R. Delany"@en 41 9.0
"Bernard Cornwell"@en 40 16.6
"Enid Blyton"@en 40 6.8
"Dr. Seuss"@en 40 26.5
The average Popularity Score that you see in the third column corresponds to the number of references (links) from other wikipedia articles to these books. Interestingly, Agatha Christie as well as Isaac Asimov are rather prolific authors whose books also have an above the average popularity. On the other hand, Ruth Rendell or Dean Koontz are rather prolific, but not very popular (at least according to wikipedia). Most popular in this list are Stephen King and Terry Prachett.

Well, let's turn it the other way around. Let's sort this list by the average popularity of the books of these authors....

Here is the Top 40 list of the authors with the most popular books (on average):
name numOfWorks popularityOfWorks
"John Simpson (lexicographer)"@en 1 1627.0
"John Milton"@en 1 669.0
"Kenneth Grahame"@en 1 423.0
"Emily Brontë"@en 1 387.0
"Wilhelm Grimm"@en 1 343.0
"Jacob Grimm"@en 1 343.0
"Harper Lee"@en 1 334.0
"Lewis Carroll"@en 6 319.8
"Miguel de Cervantes"@en 4 310.5
"Jonathan Swift"@en 2 303.5
"Wilbert Awdry"@en 1 296.0
"Cao Xueqin"@en 1 291.0
"Giovanni Boccaccio"@en 1 287.0
"William Shakespeare"@en 3 264.6
"Ian McFarlane"@en 1 255.0
"Antoine de Saint-Exupéry"@en 1 253.0
"Margaret Mitchell"@en 2 242.0
"Suetonius"@en 1 229.0
"Roger Hargreaves"@en 2 214.0
"Joseph O'Neill (writer)"@en 1 211.0
"T. S. Eliot"@en 2 186.0
"George Bernard Shaw"@en 2 184.5
"Harriet Beecher Stowe"@en 3 173.6
"Petronius"@en 1 162.0
"Charles Dickens"@en 30 159.5
"Johanna Spyri"@en 1 156.0
"Herman Melville"@en 7 154.7
"Jaroslav Hašek"@en 1 152.0
"Pierre Choderlos de Laclos"@en 1 152.0
"Oscar Wilde"@en 3 148.0
"Dave Arneson"@en 3 146.0
"Ngô Sĩ Liên"@en 1 145.0
"John Eric Holmes"@en 2 143.5
"Carlo Collodi"@en 1 142.0
"George Orwell"@en 11 135.0
"Erik Mona"@en 3 134.6
"Daniel Defoe"@en 5 132.8
"Monte Cook"@en 3 132.3
"Apuleius"@en 1 132.0
"Dan Brown"@en 6 126.6
Possibly you have never heard of John Simpson? But you will have heard about the Oxford English Dictionary. Well John Simpson was its Chief Editor...that makes sense, doesn't it? What about Kenneth Graham? Maybe you know his 1908 published novel The Wind and the Willows...

In this list it seems that it is more about literary excellency. Only one author with a rather prolific output is found, which is Charles Dickens with 30 listed Books.  But, to find also Dan Brown on this list tells me, that popularity doesn't hold for literary excellency or quality. At least he is last among the Top 40 after Herman Melville, George Orwell, Daniel Defoe, Lewis Caroll or Jonathan Swift. On the other hand, John Milton did not become rich with his one shot Paradise Lost although it is rather popular.

Here are the links to the online queries to get the most recent and complete results:
Enjoy....I'll be back, when I will find again something interesting ;-)



Friday, August 17, 2012

Who Knows Movies - The 2nd Round

Play it
We are happy to announce that our paper
Andreas Thalhammer, Magnus Knuth and Harald Sack: "Evaluating Entity Summarizations Using a Game-Based Ground Truth" 
has been accepted for the Evaluations and Experiments Track of ISWC 2012. We want to thank all of you very much, who played our WhoKnowsMovies? game, which allowed us to collect the necessary data for this publication. And since we also want provide updated statistics for the final version of our paper, you are very welcome to play a little bit more... :)

There has been a little upgrade for the game including more questions and we need more data to achieve a more reliable proof of our assumption that our proposed fact ranking for entity summarization really is better than a random choice.

Therefore, please help us and play the game. Test your knowledge about movies! Can you challenge the highscore? 

P.S. The already gathered (anonymized) data is available at http://www.yovisto.com/labs/iswc2012/.

Saturday, June 23, 2012

Who knows Movies? Another Game with a Purpose for the Semantic Web

Play it
Of course you will remember our last Quiz Game 'Who Knows' (cf. the post below), where we have used DBpedia facts to generate questionaires for a game with a purpose. The very first application of this game was to put to evaluate some newly developed entity/property heuristics via crowdsourcing. Then, we realized that the gathered data also had some collateral benefits such as the detection of inconsistencies and flaws within the Linked Data resources that we used for generating the questions.

Now, we are looking into another task: entity summarization. Entity summarization means that we are trying to wrap up only the most important facts that determine a distinct entity. Just think of the Google knowledge graph that displays entity summaries from Freebase. Thus, Entity summarization of course is rather similar to relevance ranking of facts.

To generate a ground truth for evaluation of entity summarization heuristics, we adapted our quiz game WhoKnows? to become WhoKnows?Movies! We are cooperating with Andreas Thalhammer from University of Innsbruck with this task and we have restricted the domain of questions to popular movies, adapted the gereration of questionaires by utilizing the Freebase knowledge Base. Now we have to gather data....

Therefore, please help us and play the game. Test your knowledge about movies! Can you challenge the highscore? 


P.S. Of course all the data gathered will be anonymized and made publicly avaible.

Saturday, June 25, 2011

New Challenge! Playing WhoKnows? to develop a new Teflon Pan


Some of you already may know our --serious-- fun game WhoKnows? (N. Ludwig, J. Waitelonis, M. Knuth, H. Sack: WhoKnows? - Evaluating Linked Data Heuristics with a Quiz that Cleans Up DBpedia) that has been presented inter alia at this year's ESWC 2011 (cf. picture from poster session).

WhoKnows? is a quiz game based on the DBpedia dataset; while answering the quickies the player produces data that can be used for the ranking of facts from the underlying knowledge base. Furthermore, the player has the possibility to mark strange questions often originating from inconsistencies that we want to identify this way.

Now, as a next step we want to apply the collected data for the development of an expert finder and user interest profile recommendation system. For this we would appreciate a larger data set that allows us to rate the expertise of several users in various domains. If you like to contribute this research, you can do this easily by playing WhoKnows? on Facebook. In order to make a sound statement about your expertise, we need at least about thirty questions answered and of course the more the better.

Of course all gathered data will be anonymized before analysis and evaluation.

Don't forget it's really fun and educational together!
Your help really is appreciated. Thank you for playing!

P.S. We would be pleased to inform you about the final results, if you are interested in. Just send us an e-mail.

Wednesday, November 24, 2010

'Who knows?' - A Semantic Web Game

Please support our research by playing our Semantic Web Game 'Who Knows'!

What is 'Who Knows?'
'Who Knows?' is a simple Q&A Game in the style of 'Who wants to be a Millionaire'. The questions are automatically generated from DBpedia content.

What is the purpose of 'Who Knows?'
The purpose is the evaluation of some heuristics that are used to determine a ranking of facts within a knowledge base such as e.g. DBpedia.

These are the simple assumptions 'Who Knows?' is based on:
  1. If a user knows the correct answer, the fact seems to be 'important'.
  2. If a user doesn't know the correct answer, the fact seems to be not so 'important'.
  3. If a user votes the question to be wrong, odd, or strange, the fact seems to be 'irrelevant'.
There a different variants to play the game:
  1. One-on-One questions -- only one choice is correct.
  2. N-to-One questions -- there are multiple correct answers.
  3. Hangman -- find the answer by playing the popular game of hangman.
  4. Maths -- find the answer and compute a simple arithmetic formula.
Meanwhile you will receive points for correct answers. The faster you provide the answer, the more points you will get. If you provide the wrong answer, you'll loose a life and some points will be taken from your score.

Try to score as many points as possible and don't forget to tell your friends!!!!

Monday, October 26, 2009

Open PhD Positions in Semantic Multimedia Retrieval Project

OPEN Ph.D. POSITIONS at Hasso-Plattner-Institute (HPI), Potsdam (Germany) starting on the fourth quarter of 2009

Hasso-Plattner-Institute (HPI) is a privately financed institute affiliated with the University of Potsdam, Germany. The Institute's founder and benefactor Professor Hasso Plattner, who is also co-founder and chairman of the supervisory board of SAP AG, has created an opportunity for students to experience a unique education in IT systems engineering in a professional research environment with a strong practice orientation.
(for more information on HPI, c.f. http://www.hpi.uni-potsdam.de/ )

Project Description:
MEDIAGLOBE is part of the THESEUS research program initiated by the German Federal Ministry of Economy and Technology (BMWi), with the goal of developing a new Internet-based infrastructure in order to better use and utilize the knowledge available on the Internet. The focus of the research program is on semantic technologies, which determine contents (words, images, sounds, and videos) not through conventional methods (e.g., combinations of letters) but which are able to recognize and place the meaning of a content in its proper context. MEDIAGLOBE deals with digitalization, analysis, and semantic retrieval of historical, documentary audiovisual content. (for more information on MEDIAGLOBE, c.f. http://theseus-programm.de/theseus-mittelstand-2009/ )

The ideal candidate holds a MS degree in Computer Science or related field and is able to consider both theoretical and practical/implementation aspects in her/his work. Fluent english communication and programming skills are fundamental requirements. Since we are working on a multimedia repository with resources in German language, German language skills are welcome! Preferably the candidate has a background in one of the following
fields:
• semantic web technologies
• knowledge representations and ontology engineering
• audiovisual retrieval and analysis
• semantic search
• innovative web development
• user interface design for audiovisual content

The position starts as soon as possible and is full-time (40h/week) for the duration of the project until Oct 2011. Review of applications will begin immediately and will continue until the position is filled. The successful candidate will tightly work with international partners and has the possibility to pursue PhD work within the scope of the project.

How to apply:
Excellent candidates are invited to apply with:
• Curriculum vitae and copies of degree certificates/transcripts,
• Writing samples/copies of relevant scientific papers (e.g. thesis, etc.),
• Letters of recommendation.

Please send your application in PDF format indicating in the subject 'Application for PhD position‘ via email or via traditional mail to the following contact.

Contact and application:
Harald Sack
Hasso-Plattner-Institut für Softwaresystemtechnik GmbH
Universität Potsdam
Prof.-Dr.-Helmert-Str. 2-3
D-14482 Potsdam, Germany
phone: +49 (0)331-5509-527
fax:
+49 (0)331-5509-325
email:
harald.sack@hpi.uni-potsdam.de
web:
http://www.hpi.uni-potsdam.de/meinel/persons/sack.html

Tuesday, September 15, 2009

Corporate Semantic Web Workshop in Berlin, 15.09.2009


Heute findet im Rahmen der XInnovations 2009 der Corporate Semantic Web Workshop an der HU Berlin statt, der heute abend im 3. Semantic Web Meetup seinen Abschluss finden wird.

9 Uhr morgens, noch ist alles ruhig. Kaffee allerdings scheint am morgen ein Fremdwort zu sein. Zumindest reicht es 30 Minuten vor Beginn der Veranstaltung lediglich für ein anscheinend vom Vorabend "geplündertes" Buffet mit zwei leeren Thermoskannen und ein Paar verlorener und zum Teil gebrauchter Kaffeetassen. Da lob ich mir doch Graz und die I-Semantics, bei der wir vor gut einer Woche rund um die Uhr mit leckerem illy-Kaffee (Ja! ich würde mich über eine Product-Placement-Vergütung sehr freuen...:) und Espresso versorgt wurden....

Prof. Adrian Paschke von der FU Berlin eröffnet den Corporate Semantic Web Workshop mit einer kurzen vorstellung der 'Vision' des Corporate Semantic Web Projekts, das vom BMBF seit 2008 gefördert wird. Für mich natürlich interessant der Unterpunkt und Forschungsbereich 'Corporate Semantic Search', also bin ich auf den später geplanten Vortrag zum Thema gespannt.

Gökhan Coskun aus Adrian Paschkes Forschungsgruppe schließt an mit einem Vortrag über 'Effiziente Verwaltung von Unternehmenswissen - Corporate Ontology Management'. Wozu braucht man Ontologien in einem Unternehmen? Ganz einfach, zur Steigerung der Produktivität und der Effizienz der Informationsverarbeitung (typische Wirtschaftsinformatikerantwort...). Als Hinderungsgrund für den Einsatz im Unternehmen identifiziert Coskun die 'Akademische Orientierung' vorhandener Werkzeuge. Er sieht Ontologien als normativ und allgemeingültig, was einen weiteren Hinderungsgrund bzgl. deren Einführung im Unternehmen darstelle. Dem möchte ich widersprechen, da Ontologien stets auf einer 'gemeinschaftlichen Vereinbarung' beruhen (explicit, formal specification of a shared conceptualization), die den jeweiligen Blickwinkel der Beteiligten abbildet. Allgemeingültigkeit wäre eine Eigenschaft, die gar nicht erreicht werden soll. Wir betreiben ja schließlich Informatik und nicht Metaphysik, d.h. unser Ziel ist nicht die normative und allgemeingültige Beschreibung der gesamten Welt, da eine Ontologie stets der Interpretation des Benutzers, seinem Kontext und seiner Pragmatik unterliegt.

Nach der viel zu kurz geratenen Kaffeepause geht es weiter mit Olga Streibel von der FU Berlin und dem Thema 'Semantische Suche: Tagging und Wissensgewinnung'. Als 'Extreme Tagging' werden jetzt Tags eingeführt, die selbst als Objekte für das Tagging hergenommen werden können, d.h. die eigentliche Tag-Relation lässt sich taggen. Was gewinnt man dadurch? Streibels Erklärung, dass man 'durch Tags Assoziationen bildet, mit denen man Ontologien erzeugt', hilft mir hier nicht besonders weiter. Leider wurde meine diesbezügliche Frage am Ende mit der Bemerkung, dass wir das offline diskutieren sollten, etwas unschön abgebügelt. Dabei wäre es meines Erachtens für den Vortrag von zentraler Bedeutung, eben diesen Vorteil des Ansatzes herauszustellen und gegenüber einem einfachem (individuellem) Konzept-Mapping abzugrenzen.

Ralf Heese von der FU Berlin referiert als nächstes über 'Einfach Verlinken in Wikis / Experten mittels Wikis finden'. Es geht beim ersten Thema dabei darum, den Benutzer mit Hilfe von Hintergrundwissen beim Setzen von Links durch entsprechende Vorschläge zu unterstützen. Das zweite Thema widmet sich der Frage, wie sich aus den History-Logdaten eines Wikis Experten zu bestimmten Themen bestimmen lassen.

Gleich Mittagspause....! Leider nur mit Gulaschsuppe/Möhrensuppe und Rundgang über den 'Büchermarkt' im Hof vor dem HU-Hauptgebäude.

Wieder einmal ein paar Minuten zu spät bei Richard Hubers (FIZ Chemie GmbH) Kurzvortrag über 'ChemgaPedia - virtuelle Forschungsumgebung'. Und schon wieder eine Wortneuhülse: 'Wissenscloud' - ohne sich darüber im Klaren zu sein, was exakt damit gemeint ist.

Änderung im Programm, 'Semantic Profiles in Universal Plug and Play AV', Vortrag eines Diplomanden zur semantischen Anreicherungen von UPnP AV Daten.

Weiter geht es mit Johannes Krug von x:hibit, der das Berliner Museumsportal (finanziert durch entsprechende E-Commerce Anteile, z.B. E-Ticketing, u.a.) vorstellt, gefolgt von Radoslav Oldakowski von der FU Berlin mit dem Thema 'Semantische Datenintegration und Suche im Museumsportal Berlin'. Semantisch unterstützte Suche, die Suchbegriffe um verwandte Begriffe ergänzt...mal sehen, ob sie dies auch (1) intelligent und (2) visuell ansprechend tut.

Sebastian Hellmann gibt mit 'DBPedia Live Extraction' als erstes eine kurze Einführung in das zentrale Hub der Linked Open Data Cloud, der DBPedia, gefolgt von diversen Anwendungen rund um die DBPedia. Übrigens nutzt yovisto.com (including semantic features) ebenfalls DBPedia-Daten zur Implementierung einer echten explorativen (semantischen) Suche.

[...to be continued @ Semantic Web Meetup ]

Monday, September 14, 2009

W3C-Tag an der HU Berlin, 14.09.2009

Es ist mal wieder soweit: W3C-Tag an der HU Berlin im Rahmen der XInnovations 2009, diesmal mit Prof. Felix Sasaki als Vertreter des lokalen W3C-Büros an der Uni Potsdam....und es geht auch schon gut los. Prof. Robert Tolksdorf, der die Veranstaltung eröffnen sollte, steckt im Berliner Verkehr fest und so warten wir erst einmal 20 Minuten bis hier irgendetwas heute morgen passiert.

Atemlos kommt Herr Tolksdorf mit 20 minütiger Verspätung an, entschuldigt sich kurz (ohne dabei nicht auch einen kurzen Hinweis auf die Situation der Berliner Verkehrsbetriebe zu geben und auf seine Solidarität mit den Berliner S-Bahn-Fahrern hinzuweisen) und stellt das STI (Semantic Technology Institute) vor.

Interessanter wird es jetzt schon mit dem ersten, dem 'Semantic Web' gewidmeten Vortrag, genauer geht es dem Titel entsprechend über das 'Rule Interchange Format' (RIF) und das neue OWL 2, von Prof. Adrian Paschke von der FU Berlin. OWL 2 verwirft die übliche Dreiteilung in OWL-Light, OWL DL und OWL full und definiert verschiedene OWL DL Sprachprofile bzgl. ihrer 'worst case' Berechnungskomplexität. Dabei lässt sich OWL EL sogar in Polynomialzeit berechnen (daneben existieren noch OWL QL und OWL RL). RIF als Austauschformat für Regeldialekte bringt ebenfalls wieder eine Unmenge an neuen Syntaxvarianten. In diesem Zusammenhang wird auch auf ein Handbuch hingewiesen (Handbook of Research on Emerging Rule-Based Languages and Technologies, IGI Global), das sich allerdings nicht gerade durch seinen Preis (> 300 Euro) empfiehlt...

Prof. Felix Sasaki vom deutsch-österreichischen W3C-Office stellt als nächstes die aktuellen Entwicklungen rund um den neuen HTML5 Standard vor (hier ein Link auf die im Vortrag gezeigten Beispiele). Warum eigentlich jetzt HTML 5, nachdem bereits 2000 XHTML 1.0 veröffentlicht wurde und die Entwicklung von XHTML 2.0 auf vollen Touren lief? Nun, die Entwicklung von XHTML 2.0 wurde abgebrochen, der Anspruch ein 'sortenreines' XML im Browser einzuführen (eine 'Revolution') ist gescheitert. Nach einer von Opera durchgeführten Studie 2008, ist lediglich 4.13% des gesamten Web-Codes tatsächlich valide. Daher versucht man mit HTML 5 eine (fehlertolerante) 'Evolution' des alten Standards zu realisieren.
Eigentlich ist HTML5 genau genommen sogar ein Rückschritt, da es auf eine dezentrale Erweiterung über einzubindende Namensräume zugunsten eines eindeutig zu interpretierenden DOM-Baumes verzichtet. Dahingehend sind Probleme mit semantischen Erweiterungen, wie z.B. RDFa oder Microformats vorprogrammiert!

Mittagspause im Café Chagall mit Bliny und saurer Sahne mit anschließendem Rundgang zu Dussmanns Kulturtempel...

Thomas Caspers spricht zunächst einmal über Barrierefreiheit (Die deutsche Übersetzung der WCAG 2.0 a.k.a. Web Content Accessibility Guidelines). 'POUR' steht für die vier Grundprinzipien der Richtlinien für barrierefreie Webanwendungen: Perceivable, Operable, Understandable und Robust. Allerdings konnte ich die beschriebenen Übersetzungsprobleme ('programmatically determined' usw.) nicht nachvollziehen...vielleicht hätte man mal einen Informatiker fragen sollen....

Mit gut 20 minütiger Verspätung (dank der Ausdauer des Vorredners) beginnt der Vortrag von Joachim Neuberth über SKOS (Simple Knowledge Organization Systems). Warum müssen manche Vortragende nur immer so leise sprechen. Das Verfolgen des Vortrags gestaltet sich nicht wirklich einfach (was nicht etwa in der Komplexität des Themas begründet liegt). Zugegebenermaßen ist aber auch die Entwicklung von SKOS und das dahinter liegenden Datenmodell nicht so besonders spannend. Besser wird es erst, als unterschiedliche Ressourcen, wie z.B. die Library of Congress Headings oder die französische Nationalbibliothek aufgezeigt werden und am Ende dann doch Linked Open Data angesprochen wird.

Prof. Felix Sasaki widmet sich als nächstes dem Thema Metadaten für Multimedia und berichtet von der W3C Metadata Annotations Group. Das Ziel der Metadata Annotation Group kann als "DublinCore + X" paraphrasiert werden, also ein minimales Multimediametadatenschemas mit Mapping zu bereits existierenden Formaten. Der zweite Teil des Vortrags beschäftigt sich mit XProc (XML Pipeline Language) zur Modellierung und Beschreibung von XML-Verarbeitungsketten, die eine Pipeline-artige Verarbeitung von wechselnden Validierungs- und Transformationsschritten unter Ausnutzung einer rudimentären Programmlogik (konditionale Verarbeitung, Iterationen, Selektive Verarbeitung, Ausnahmebehandlung, u.a.) mit heterogenen XML-Daten erlaubt.

[Nach der noch folgenden Diskussion, weiter mit "Brezeln und Wein" im Foyer..... :)]

Links:

Friday, September 04, 2009

i-Semantics 2009 in Graz (Day 03)

The second day of i-Semantics ended with party and dancing to live music performed by 'Egon 7', and of course with a lot of interesting talks with interesting and nice people :) This morning at breakfast, Jörg really looked as if he had not really had got enough sleep (he continued to party after the official ending somewhere downtown :)

Keynote
Peter Kropsch (Austrian Press Agency): When technologies are drivers, integrated concepts are needed for success,
talking about scenarios of future media convergence, the development of the information technologies and the possibilities opened up by them, esp. about the expectations of APA what to get out of semantic technologies. In general, keynotes without some slides (to get hold of the information structure of this 60 minutes talking) are rather difficult to follow, if the speaker is not able to awake sufficient enthusiasm in the audiences...

The Role of Semantic Technologies in Future Internet Track,
Klaus Tochtermann (Know-Center, TU Graz): The Role of Semantic Technologies in the Future Internet
explained, why the vision of the future internet is not only 'old wine in some new bottles'. Although all the components that consttute the Future Internet (Content, Services, Security, People, etc.) is already around, the main contribution of FI is the integrated view and interaction of thes components.

Marco Pistore (FBK): Highlights of the Future Internet Conference Berlin
In parallel to i-Know/i-Semantics in Berlin the 2nd Future Internet Symposium took place...

Jan Reichelt (Mendeley): Mendeley - A Last.fm for Research?
Mendeley wants to help researchers to manage their resources. Research is inherently social. Mendeley offers a desktop tool, similar to last.fm audioscrobbler that (a) analyzes your search papers and enables full-text search and (2) smart metadata extraction and generation. The mendeley web service offers an online backup of your own research library and enables on-click import from google scholar and other citation services.
Interesting stuff...I have just registered and download my desktop client for Mac OSX...

Paolo Rosso (Universidad Politecnica de Valencia): Geographical Information Retrieval and Toponym Disambiguation
Geographical filtering of information retrieval results by expanding query strings with semantically related geographical information (...like that stuff!).