Showing posts with label wikipedia. Show all posts
Showing posts with label wikipedia. Show all posts

Saturday, June 21, 2014

Harald's Original Miscellany - Prolificacy vs. Popularity in Literature, Part 2

It might have been a surprise for you that according to Wikipedia, editor in chief of the Oxford English Dictionary John Simpson is the most popular author [1]. Thus, we have to take a look on the notion of "popularity". In scientific publishing, an "important" author is an author whose works are cited (referenced) by many other authors. In this way, a ranking among the most important authors has been established. In Wikipedia, we can follow this approach and simply count, how many articles are referencing the article of an author or the articles of the author`s books. In our approach, this was simply the number of incoming pageLinks of the referenced articles. Currently, we are adopting the PageRank algorithm for achieving a better measure of "importance" of an article. You all know the PageRank algorithm named after Larry Page, one of the founders of Google. PageRank is a way of measuring the importance of website pages [2]. We will get back on this, when we have finished PageRank integration into DBpedia.

But, let's discuss the results we had achieved in our last blog post. We stated that there are 15,328 authors [3] referenced in DBpedia. Well, of course there might be more authors, but this is the number of individuals that belong to the DBpedia class authors. The question is, when we made the popularity vs prolificacy statistics, did we include the works of 15,328 authors? Surely not. Let's first check, for how many authors, there are books written by these authors referenced in DBpedia: Its only 3,763 authors with referenced books [4]. Thus, we don't have any idea about the remaining 11,000 authors who don't have one of their books referenced in DBpedia. They must have written some book, otherwise they would not be an "author"....

OK, so we are dealing with insufficient information. If we look at the available data, we could also use another property, called dbpedia-owl:notableWork denoting only the "important" works of an author (if we relate it with authors and not with other artists). Now it would be interesting, to repeat our statistics based on this property and look, if there is a significant difference. First let's look at the most prolific authors with respect to referenced "notable works":

name numOfNotableWorks popularityOfNotableWorks
"Roald Dahl"@en 16 53.8
"Joseph Conrad"@en 9 68.2
"Charles Dickens"@en 8 398.7
"Roger Zelazny"@en 7 15.7
"Henryk Sienkiewicz"@en 6 53.8
"Tom Wolfe"@en 6 44.3
"Margaret Atwood"@en 6 41.8
"Henry James"@en 6 74.0
"Ismail Kadare"@en 6 7.8
"Alfred Döblin"@en 6 13.3
"Stephen King"@en 6 133.1
"Robert Girardi"@en 6 3.0
"Hunter S. Thompson"@en 5 56.4
"Joanne Harris"@en 5 16.4
"David Mitchell (author)"@en 5 25.6
"Chris Kuzneski"@en 5 7.0
"Jackie French"@en 5 8.2
"Sita Ram Goel"@en 5 7.0
"Peter Hitchens"@en 5 9.2
"Chinua Achebe"@en 5 31.6
"Colm Tóibín"@en 5 17.4
"E. L. Doctorow"@en 5 26.4
"Don DeLillo"@en 5 27.4
"Johann Wolfgang von Goethe"@en 5 72.6
"Samuel Beckett"@en 5 25.4
"Cormac McCarthy"@en 5 55.6
"Brian O'Nolan"@en 5 29.8
"Poppy Z. Brite"@en 5 11.2
"William Trevor"@en 5 20.2
"Neil Gaiman"@en 5 84.8
"Dean Koontz"@en 5 10.4
"George MacDonald"@en 5 21.6
"Ann Bannon"@en 5 9.0
"William Faulkner"@en 4 70.7
"Will Self"@en 4 10.0
"Alan Hollinghurst"@en 4 23.5
"George Eliot"@en 4 81.0
"Malcolm Gladwell"@en 4 39.0
"C. S. Lewis"@en 4 43.2
"Hermann Hesse"@en 4 53.2

There is a huge difference. Less science fiction authors, more international authors, and the average "popularity" of the top 40 author's works has also significantly increased. We also realize that considering Charles Dickens, the popularity of his 8 notable works (398) doesn't differ so much from the popularity of his overall 30 works (159.5) as for Stephen King, where the popularity of his 5 notable works (133) is much higher compared to the overall average from last time (44) for his 75 books. Interesting also that Road Dahl is mentioned with 16 notable works. Either the wikipedia authors could not decide which of the works of Road Dahl really was notable or the article must have been written by a huge fan. OK, so we might learn that is in the eye of the beholder, which works of an author are referenced as being "notable".

Let's order the list now again according to the popularity measure and see what happens. This is the Top 40 list of authors with the most popular works (on average) with respect to "notable works" only:

name numOfNotableWorks popularityOfNotableWorks
"Bram Stoker"@en 1 1145.0
"Lewis Carroll"@en 2 939.0
"Lucifer Chu"@en 1 612.0
"Miguel de Cervantes"@en 2 611.0
"J. R. R. Tolkien"@en 2 566.0
"Robert Louis Stevenson"@en 3 448.0
"Alexandre Dumas"@en 2 427.5
"Oscar Wilde"@en 1 427.0
"Kenneth Grahame"@en 1 423.0
"Walter Prescott Webb"@en 1 423.0
"James Vincent Murphy"@en 1 421.0
"George Orwell"@en 3 412.3
"Charles Dickens"@en 8 398.7
"Emily Brontë"@en 1 387.0
"Leo Tolstoy"@en 2 374.5
"John Bunyan"@en 1 341.0
"Louisa May Alcott"@en 1 314.0
"Mark Twain"@en 2 307.5
"Jonathan Swift"@en 2 303.5
"Emma Orczy"@en 1 282.0
"Charlotte Brontë"@en 2 258.0
"Ian McFarlane"@en 1 255.0
"William Golding"@en 1 251.0
"Anne Frank"@en 1 246.0
"James Fenimore Cooper"@en 1 246.0
"Margaret Mitchell"@en 2 242.0
"Frans Sammut"@en 2 233.5
"Rudyard Kipling"@en 2 230.5
"Dan Brown"@en 3 224.3
"John Steinbeck"@en 3 218.0
"William Gibson"@en 1 216.0
"Ayn Rand"@en 2 210.5
"Harriet McDougal"@en 1 207.0
"Jim Butcher"@en 1 200.0
"Richard Adams"@en 1 192.0
"William Makepeace Thackeray"@en 1 191.0
"T. S. Eliot"@en 2 186.0
"Aldous Huxley"@en 3 175.3
"Hitoshi Igarashi"@en 1 175.0
"Alice Walker"@en 1 174.0

This looks a lot more like we would have expect it to be in the first place. Bram Stoker on the first place with only 1 notable work, we all know which book is meant by that :) impressive score. Lewis Carroll on the second place with his Alice in Wonderland stories is also no wonder. But, who is Lucifer Chu? This is some kind of surprise for me. You won't find Lucifer Chu in the German version of DBpedia, so why does our result suggest that he is a popular author? Well, if you look at the very short Wikipedia article in the English version, you will find out that he is the author of the Chinese translation of J.R.R.Tolkien's The Hobbit and The Lord of the Rings. Tolkien's books follow on rank number 5 with a slightly less popularity. How can that be? Well we must examine the available data a little bit closer.

Chu is referenced in our list with 1 book only although we know that he has translated The Hobbit as well as The Lord of the Rings. And J.R.R. Tolkien is referenced with 2 books. Obviously, for Chu only the Hobbit as being the most popular of these books has been counted. But, if we look at the DBpedia page of Lucifer Chu, we realize that for Chu there are 2 books referenced as notable Work, as there are 3 books for Tolkien (including The Silmarillion). The reason for this is that The Lord of the Rings is not a member of the class dbpedia-owl:Book. Why is that so? Maybe because The Lord of the Rings consists out of 3 different books. Thus, to get a more complete result, we should think of how we get all notable works into that list. But, we should simultaneously take care of only including books or series of books and to exclude other pieces or work such as e.g. paintings, sculptures, movies, photographs, etc.

We will take care of this next time... :)

References:
[1] Harald's Original Miscellany - Prolificacy vs. Popularity in Literature, Part 1, 2014/06/20
[2] Brin, S.; Page, L. (1998). "The anatomy of a large-scale hypertextual Web search engine". Computer Networks and ISDN Systems 30: 107–117.
[3] live SPARQL query for the number or authors
[4] live SPARQL query for the number of distinct authors who have written a book, referenced in DBpedia
[5] live SPARQL query for the Top 40 most prolific authors (based on notable works)
[6] live SPARQL query for the Top 40 most popular authors (based on notable works)

Tuesday, July 27, 2010

There are more Things in Heaven and Earth... - DBPedia Link Graph Analysis Revisited

In the course of our ongoing work with Linked Open Data, we recently made some analysis on the graph structure of DBPedia data. For this we only took under consideration the original link graph (aka 'wikilinks'), where we did some cleanup first, such as, e.g., resolving redirects, etc.

As a side effect, we had to compute in-degree and out-degree of all DBPedia entities according to wikilinks, ... and we discovered some more or less surprising facts (thanks to Nadine):

The entity with the hightest out-degree (i.e. number of outgoing links) currently is:
http://dbpedia.org/resource/List_of_places_in_Afghanistan
with 7.147 outlinks (after cleanup, and remember it's wikilinks and not typed links of DBPedia)

The entity with the highest in-degree (i.e. number of incoming links) currently is:
with 440.151 inlinks (after cleanup)

While the 2nd one (living people) seemed pretty clear to me, the first (Afghanistan places...) was a bit of a surprise (as also are trilobytes...). For all the explorers among us, I have included the Top Ten list of incoming and outgoing wikilinks, each with indegree and associated outdegree...


Top Ten Incoming in out
http://dbpedia.org/resource/Living_people 440151 0
http://dbpedia.org/resource/United_States 385407 963
http://dbpedia.org/resource/France 124206 759
http://dbpedia.org/resource/England 123223 1320
http://dbpedia.org/resource/United_Kingdom 121203 1152
http://dbpedia.org/resource/List_of_sovereign_states 114086 465
http://dbpedia.org/resource/Canada 105849 523
http://dbpedia.org/resource/Germany 103382 889
http://dbpedia.org/resource/Animal 98680 236
http://dbpedia.org/resource/World_War_II 93555 771
http://dbpedia.org/resource/Association_football 90673 196



Top Ten Outgoinginout
http://dbpedia.org/resource/list_of_places_in_afghanistan97147
http://dbpedia.org/resource/Flora_of_New_South_Wales 917 6819
http://dbpedia.org/resource/List_of_municipalities_of_Brazil 1 5503
http://dbpedia.org/resource/Index_of_India-related_articles 4 5369
http://dbpedia.org/resource/Area_codes_in_Germany 6 5360
http://dbpedia.org/resource/IUCN_Red_List_vulnerable_species_%28Plantae%29 0 5172
http://dbpedia.org/resource/List_of_trilobites6 5102
http://dbpedia.org/resource/List_of_Social_Democratic_Party_of_Germany_members 24 5078
http://dbpedia.org/resource/List_of_French_words_of_Germanic_origin 9 5010
http://dbpedia.org/resource/Index_of_Thailand-related_articles 4 4831

But there are more interesting things to discover ... stay tuned!

Monday, September 22, 2008

XInnovations 2008, Berlin, Day 01, Sept. 22, 2008

For the third time, I'm attending the XInnovations (formerly known as XML-Days) 2008 in Berlin. Although I was often rather dissapointed about the quality of the conference program (you might refer to my previous posts about XML-Tage Berlin here), I decided to give it another try (simply because I can reach Humboldt University in Berlin with local public transportation in about 45 minutes). Also this time, there seemes to be some emphasis on semantic web technology (at least considering the program, there are Semantic Wikis in the Corporate Wiki track and anoteher Corporate Semantic Web Workshop, not to forget the Semantic Web topic in the PhD-Forum).

I started the first day of the coference with participating the Coprporate Wiki Infotag. Denny Vrandecic is talking about the "Semantic Media Wiki" from AIFB Karlsruhe. Denny is starting his talk with some historical facts about WIkipedia. Interesting thing to mention, according to a study from Aaron Swartz, only 2% of all Wikipedia users (it come up to 1.400 people) are primarily responsible for all article changes. This contradicts the commonly assumed opinion that wikipedia is written by millions of users. Then he was introducing the semantic web in general, by stating that the semantic web is nothing but things (nodes, concepts) being connected by certain relationships, forming graphlike structures that can again be related to each other. According to his definition, a semantic wiki is nothing but graphs being created from wiki data.

The next talk I'm attending is in the Ph.D workshop. Olaf Hartig is talking about 'Trustworthiness of Data on the Web'. With the Semantic Web more and more software agents are taking decisions based on (RDF-based) data on the web. But how can we trust those data? Olaf is developing an RDF trust model as a basis for trust assesment and trust-aware data acces. He suggests a scale from [1-;1], where -1 represents 'absolute distrust' and +1 'absolute trust' for a statement. Now, all relationships in an RDF-graph can be weighted with according trust values ranging from [-1] to [+1]. Trust into a set of statements can be expressed with aggregated trust functions ranging from a cautious (conservative) minimum to a slightly optimistic median. The formal trust vocabulary can be found here. Next, criteria for trust assessment are collected and three different trust assessment strategies are defined: user-based (ask the user on his/her opinion about the trustworthiness), provenance-based (taking into account the trustworthiness of the referring users), and opinion-based (recommendations by other users according to their own trustworthiness).

The afternoon session started with Nils Barnickel from Fraunhofer IOCS with a talk on 'Semantic Mediation between Loosely-Coupled Information Models in Service Oriented Architectures'. Semantic descriptions of Web Services are supposed to enable data and service interoperability. One problem being addresses ist the lack of efficient ontology mapping options in current existing onlology languages (although OWL does have a differentFrom or sameAs operator, complex mappings deploying concepts with totally different subgraphs or 1:n, n:m mappings are missing).

Ok, now it's definitely time for a coffee break. after that, I will be joining the 'World Cafe' session, where I will participate in the discussions instead of writing blog.

[to be continued tomorrow, XInnovations Day 02, Corporate Semantic Web Workshop]

Thursday, September 27, 2007

CSSW 2007 - Conf. on Social Semantic Web in Leipzig, Sep. 27th 2007 - Day 02

The second day of the Social Semantic Web Conference (CSSW 2007 here in Leipzig. Fortunately, Weimar has a rather good train connection to Leipzig (approx. 55 minutes...). Thus, I can sleep at home and don't have to stay in a hotel in Leipzig (and ofcourse the costs for travelling will be reduced...as we have a very low budget at the University for conference travells). Ok, first thing I had to learn was, the conference catering isn't really as bad as I had written yesterday. Actually, you get 5 tickets for refreshments every day (not for the whole conference). Thus, you won't die of thirst ;-)
(the Picture above is showing the University of Leipzig, Jahn Campus, where CSSW 2007 together with SABRE07 takes place)

But, back to the conference program. Today, I will chair the first session and therefore, I'm not able to blog live (at least during the morning). But I will write about the talks as soon as I find some spare time. The following talks are scheduled: Patrick Maué from the University of Münster with the topic 'Collaborative Metadata for Geographic Information'. Patrick is addressing so called participatory Geographic Information Systems (GIS) that are used for decision making, as e.g. in urban planning. Usually this data is published as a catalogue on the web. But, queries to this catalogue suffer of very bad recall and precision. This also comes from the dicersity of people involved in the process of generating metadata. People have different perspective, mental models, and terminology. Thus, the general problem being addressed is refinement of shared metadata. Of the three diffeent levels of semantics ( implicit semantics whic means data gathered from statistical analysis of the original data, soft semantics such as e.g. folksonomies, and formal semantics that allow proper reasoning), Patrick addresses the gap between implicit and soft semantics, and tries to bridge this gap by detecting similarities of metadata.

The following talk by Thomas Riechert has the title'Mapping Cognitive Models to Social Spaces - Lightweight Collaborative Development of Project Ontologies'. There, an application of the SoftWiki ontology for requirement analysis (SWORE) in the software development process is presented. Software development is carried out collaboratively by the different stakeholders of the process with the goal to model an application on an abstract level from different points of view. Stakeholders formulate their requirements in natural language and use the SoftWiki for coordination and aggregation of the requirement model, the project model, and the bug model. Then, stakeholders tag the requirements and from the tags the system extracts relationships between requirements that finally end up in the project model (=project ontology). The content of the three models serves as input to the Software Development process (a.k.a. CASE-Tool), which maps the models to UML and in this way creates the input data for the software developers.

Sören Auer (also co-chair of the conference) concludes the session with a talk on'DBpedia Relationship Finder'. Firt at al, DBpedia is an interesting project that makes use of the (inherent) structured data in wikipedia. This structured data can be found in so called 'info-boxes', i.e. tables usually put in a columns right of the wikipedia article containing data in a well structured (and hopefully commonly agreed) format. This data can be extracted from the wikipedia dump. The structured data is transformed into RDF-Triples that constitute a huge graph. The goal of DBpedia is (in the end) to enable the user to ask complex queries on the structured wikipedia data. The general problem anyway is to visualize this huge amount of data in an efficient way. The relationship finder tries to draw connections between two terms in DBpedia and therefore traverses the DBpedia graph trying to find paths (via different properties being specified by extracted content the original wikipedia info-boxes). These paths are presented in shortest path first order. Up to now, no ranking of paths with similar length is performed. Interesting thing to mention ist that the terms 'Leipzig' and 'Semantic Web' are connected via the property of Leipzig being the city, where Johann Sebastian Bach died ;-)
To get more information, you can exclude certain properties from the result paths (i.d. then only paths, which don't include that specific property are included).

The upcoming session is focusses on presentations around the SoftWiki project (as also was the talk of Thomas Riechert in the first session). The following talks are scheduled: Kim Lauenroth is presenting 'A Processmodel for Wiki-based Requirements Engineering Supported by Semantic Web Technologies', followed by Haiko Cyriaks with 'Supporting Requirements Elicitation by Semantic Preprocessing of Document Collections', followed by Steffen Lohmann with 'Ways of Participation and Development of Shared Understanding in Distributed Requirements Engineering', and Thomas Riechert with 'Towards Semantic Based Requirements Engineering'.

The Session after the lunch break starts with Stefan Kröger from the University of Potsdam with 'Analysing Wiki-based Networks with SONIVIS. The main goal of the Project SONIVIS is (or at least one of the goals...) the Unterstanding of Emergence of Knowledge in Social Knowledge Spaces. As a tool, SONIVIS integrates analysis, evaluation, visualization, and data handling of wiki-based networks.
Rainer Hammwöhner from the University of Regensburg is next talking about 'Semantic Wikipedia -- Checking the Premisses'. For this reason, they took samples from (different multilingual versions of) wikipedia and tested, as e.g., if the wikipedia category system really is a sound taxonomy (...I would say 'no'!). As we already have thought, the category system makes inadequate use of hierarchies, and that the quality of different language versions varies. There seems to be a high amount of disagreement in the category systems of the different language versions.
Joshua Bacher from the Max Planck Institute for Evolutionary Anthropology is giving a Demo Talk on 'BoWiki' - a collaborative editor for biomedical ontologies, gene functions and annotations (originally derived from Semantic MediaWiki, but for being able to use reasoning SMW was given up).
The next demo is given by Marc Fleischmann entitled sMeet-Let's Meet real, a web platform to talk and to sozialice (in an synchronous way) just as in real life. sMeet constitutes a 3D Avatar based virtual community (just as 2nd Live), but connects the virtual world with the phone system....(really an entertaining presentation, I even saw my very first live iPhone...But, I`m missing semantics....). The phone system enables real mobility (a kind of ambient 2nd Live....and the phone system is also something that people are used to pay for) and with sMeet several heterogeneous communities are (at least planned) to be connected.
Richard Cyganiak closes the session with a presentation on 'DBpedia - a Nucleus for a Web of Open Data' providing more background information on the DBpedia project. In DBpedia, every item has an own URI. Simply take the wikipedia URI of an article and substitute 'www.wikipedia.org/wiki' by 'www.dbpedia.org/resource' (here you might find the DBpedia resource named 'Leipzig'). Thus, DBpedia becomes a repository for Semantic Web identifiers.

The final event of today (before the conference diner) was a panel discussion on the topic 'Is there a Social Semantic Web?', in which I took part (therefore no live-blogging). I will refer to that in my next post....

Thursday, May 03, 2007

A map of the Social Web


Via media-ocean I found an interesting map of the 'currently known' Social Web (a.k.a. Web 2.0). Here you might take a look at the map in full size. Interesting thing about, the sizes of the single areas (communities) on the map correspond to the approximate number of members. Also the layout has its purpose. On top, you'll find the 'practicals' (Yahoo, Windows Live...), while at the bottom there stick the 'intellectuals' (wikipedia, sourceforge,...). The left is more concerned about 'real life' (I'm missing XING...), while the right is more 'web centered' (second life, antropomorphic dragons, etc...). You can even find good old usenet ...but only as a dashed outline (maybe it is about (or has already) to vanishid and only remembered in old myths like the legendary Atlantis..)
Nice thing to notice, there's also Qwghlm....I can't remember having seen it on any map before :)

Saturday, January 13, 2007

...against all odds


On wednesday I attended a talk given by Michael Strube from EML Research on "World Knowledge induced from Wikipedia - A New Prospect of Knowledge-Based NLP ". He was showing how the (meanwhile famous) collaborative encyclopedia can be used for information retrieval purposes in a way similar to (more traditional) online dictionaries as e.g. WordNet and - though being not well structured - provides results of almost equal quality.
First thing was that for their work, Strube and his colleague regarded each Wikipedia page as being the representation of a concept (we already had some arguments about that as you might remember...). Next, they developed some metric for similarity of concepts w.r.t. to the concept hierarchy (where the wikipedia defined 'concepts' come into play). Since 2004, wikipedia features a user defined concept hierarchy. This hierarchy of concepts also can be regarded as being a folksonomy, simply because this is not a knowledge representation carefully designed by some designated domain expert, but by the wikipedia comunity in a collaborative way. Unfortunately, the wikipedia concept hierarchy suffers exactly from that fact. From my pont of view it seems problematic to compare the proposed similarity measure (based on wikipedia concept hierarchy) with other similarity measures (based on commonly shared expert ontologies). O.k., you might argue that indeed the wikipedia concept hierarchy IS commonly shared, because it has been developed by the wikipedia community...but is the knowledge represented in wikipedia really 'common'? Just remember the diversity and manifold of Star Wars characters or Star Trek episodes in wikipedia compared with, as e.g., the history of smaller Eropean countries. As for all ontologies always the view and the knowledge of the ontology designer has to be considered. The wikipedia concept hierarchy - although partly being really appropriate - reminds me somehow to this famous literary chinese dictonary entry defining the term 'animal' which is quoted by Jorge Luis Borges. Another problem lies in the fact that the different language versions of wikipedia have developed different concept hierarchies (sic!).

In the end, I was asking how this proposed information retrieval based on wikipedia could be improved by considering a 'Semantic Wikipedia', as e.g., the Semantic MediaWiki (given that those semantic wikipedias would contain sufficient data). Instead of answering my question, Michael Strube cited Peter Norwig's argument against the Semantic Web from last years AAI2006. Just to sum up: the semantic web will not become reality because of the inability of its users to provide correct semantic annotations. But hey...this guy (Strube) was talking about wikipedia. Doesn't this argument raise any associations? Just remember the time 5 or 10 years ago. Nobody (well almost nobody) would have believed that it will be possible to write an entire encyclopedia collaboratively on an open source basis - just because the web user's did not seem to be able to write 'correct' articles....

Wednesday, November 15, 2006

wikipedia to serve as a global ontology....



Today, I met Lars Zapf for a quick coffee enjoying the rare late afternoon november sun. We were exchanging news about ISWC, WebModay, recent projects, and stuff like that. While talking about semantic annotation, Lars pointed out that instead of using (or developing) own ontologies for annotating (and authoring) documents, you could also use a wikipedia reference to indicate the semantic concept that you are writing about. Thus, as he already wrote in a comment, e.g., you could use the link http://en.wikipedia.org/wiki/Rome to indicate that you are refering to the city of Rome, the capital of Italy.
Of course you might object that there are several language versions of wikipedia and thus, there are several (different) articles that refer to the city of Rome. To use wikipedia as a 'commonly agreed and shared conceptualization' - to fulfill at least some points of Tom Gruber's ontology definition as long as wikipedia lacks the 'formal' aspect of machine understandability - we can make use of the fact that articles in wikipedia can be identified with articles in other language versions with the help of the language indicators at the lower left side of wikipedia's user interface. To serve as a real ontology, each wikipedia article should (at least) be connected to formalized concept (maybe encoded in RDF or OWL). This concept does not necessarely have to reflect all the aspects that are reported in the natural language wikipedia article. E.g., Semantic Media Wiki is working on a wiki extension to capture simple conceptualizations (such as e.g. classes or relationships).
An application for authoring documents could easily be upgraded by offering links to related wikipedia articles. If the author enters the string 'Rome', the application could offer the related wikipedia link to Rome [or any selection of related offers] and according to the authors this link can be automatically encoded as a semantic annotation (link).
O.k., that sounds pretty simply. Are any students out there to implement it (anybody in need for credit points??)? I would highly appreciate that...