Showing posts with label ontologies. Show all posts
Showing posts with label ontologies. Show all posts

Saturday, September 15, 2007

Taxonomy .... really it's a battleground


Taxonomy as a science has been founded by the 18th century Swedish botanist Carl Linné (later enobled Carl von Linné or more fashionable in Latin Carolus Linnaeus). He was born in 1707 (we are celebrating his 300th birthday!!) and after some difficulties to start (he was a rather sluggish student and his dissapointed father saw no other option than to apprentice him to a cobbler...Linné soon realized that academia might not be the worst choice and begged for a second chance, which was granted) he studied medicine in Sweden and Holland. But, his passion should become nature and all living things. He started to write catalogues of the world's plants and animal species, using a system devised by his own, which lead to his great Systema Naturae and made him famous. He is reported not being the most modest man of his time. He even suggested that his grave stone should bear the inscription Princeps Botanicorum (similar to the title having been granted to Carl Friedrich Gauss, the Prince of Mathematicians).

Linné classified all living things on earth according to its physical attributes. The idea is to categorize everything hierarchically. A certain species belongs to a special genus, several genera belong to a special family, which is further summarized in orders, classes, subphyla, phyla, kingdoms, and domains. So, e.g., man is of the genus Homo and of the species Sapiens. We belong to the family of Hominidae, which is part of the order Primates, which belong to the class Mammalia, which belong to the subphylum Vertebrae, belonging to the phylum Chordata. Furthermore we belong to the kingdom animalia and the domain eucaria. This is the Taxonomy we use today. At the times of Linné genus and species were already in use for about a hundred years, the term phylum (being adapted from the Greek φυλαί [phylai], the clan-based voting groups in Greek city-states) has first been coined by one of our local heroes (remember, I'm working at the University of Jena) Ernst Haeckel in 1876.

Although the system of taxonomy seems to be pretty straightforward, the trouble is to determine how many categories to divise from each section above. Even at a such basic level as phylum, which determines the very basic building plans of organisms, Wikipedia refers to 35 phyla, other biologists opt for a total of about thirty, others for even less. The American entomologist Edward O. Wilson actually votes for 89 phyla. So, where to decide the divisions...?
The other problem is double entries. There are, e.g., more than 5000 species of grass -- all of them looking rather similar. Some of them are reported to being inserted in taxonomy under about twenty different names independently by different scientists.

So, the general trouble seems to be common agreement. Common agreement is a fundamental principle of ontologies -- and taxonomies are a form of lightweight ontologies. There are millions of species populating our planet. Taxonomy is the prime tool to put (some) order in this diversity. Linné introduced his system in the early 18th century. For more than 200 years now biologists are trying to work on an unambiguous and consistent classification scheme and try to populate it a consistent way. Still today they have not finished. Besides some quarrels there is a vast ammount of species that has not been discovered now. So, the work being invested into the science taxonomy will continue...and who knows if it will ever end....
Thinking of ontologies and the Semantic Web, I think we are in some similar situation as biology has experienced during the last 200 years. Computer scientists are quarrelling about top level ontologies and there are many independently developed (overlapping and ambiguous) domain ontologies. Ontological Engineering -- including ontology mapping, ontology alignment, and ontology merging -- tries to find some way to get along even with inconsistencies...

BTW, ontology mapping then should also find a way to map Linné's taxonomy to Jorge Luis Borges "The Analytical Language of John Wilkins," where he describes 'a certain Chinese Encyclopedia,' the Celestial Emporium of Benevolent Knowledge, in which it is written that animals are divided into:
(1) those that belong to the Emperor,
(2) embalmed ones,
(3) those that are trained,
(4) suckling pigs,
(5) mermaids,
(6) fabulous ones,
(7) stray dogs,
(8) those included in the present classification,
(9) those that tremble as if they were mad,
(10) innumerable ones,
(11) those drawn with a very fine camelhair brush,
(12) others,
(13) those that have just broken a flower vase,
(14) those that from a long way off look like flies.



More on Carl Linné, taxonomy in biology, and 'nearly everything' you may find here:
- Bill Bryson: A Short History of Nearly Everything, Doubleday, 2003.

Monday, June 04, 2007

ESWC 2007 - European Semantic Web Conference, Innsbruck (Day 01)



Today, the 4th European Semantic Web Conference starts in Innsbruck (Austria). I arrived already yesterday in the evening by train. By chance, I already met 2 other participants (Pascal Hitzler and Andreas Hotho ... -> bibsonomy) while waiting at Munich train station. Thus, the last part of the train ride was quite entertaining :)

Keynote 1 - 9.00 -10.00
Day 1 of ESWC 2007 is starting with a keynote of Stefano Ceri on 'Design Abstractions for Innovative Web Applications: the case of the SOA augmented with Semantics'. Sorry to say, but -- at least for me -- the talk was not really interesting. Since I am not so much interested in Semantic Web Services, I decided to spent the rest of the morning sessions (after the coffee break) in the Ontology Learning, Inference and Mapping session. In parallel there are session on Semantic Web Services and Semantic Web Use Cases (unfortunately in areas also not very interesting for me, i.e. mechatronics and e-Governement). But, I'm looking forward to the Ontology Engineering session....

Ontology Engineering I - 10.30 - 12.30
The Ontology Engineering session starts with Martin Hepp on 'GenTax: A Generic Methodology for Deriving OWL and RDF-S Ontologies from Hierarchical Classifications, Thesauri, and Inconsistent Taxonomies', where he is referring to a novel methodology for automatically deriving consistent RDF-S and OWL ontologies from informal hierarchies. He demonstrated the usefullness of the approach by utilizing it for transforming the two e-business categorization standards eCl@ss and UNSPSC into ontologies. The approach is based on drawing a random sample from an input classification, letting a human user confirm/reject/adapt the sample, and using this confirmed sample as a basis for further automated classification. The tool will be available here soon.

The second talk is presented by Maciej Janik on 'SPARQLeR: Extended Sparql for Semantic Association Discovery'. Semantic Associations (i.e. complex semantic relationships among entities) in knowledge bases have to be discovered, which comes to the process of finding paths of possibly unknown length that connect the given entities and have a specific semantics. SPARQL for querying RDF databses is extended for the discovery of such paths including the possibility of formulating regular expressions over properties for specifying the required semantics of the queried paths.

Eyal Oren from DERI Galway is speaking about 'Algorithms for Predicate Suggestions using Similarity and Co-Occurrence'. Only through shared vocabularies can meaning be established. Tagging Systems solve this problem by tag suggestions to achieve coherrent vocabulary. The same principle holds for creating distributed RDF databases. Therefore, two domain-independent algorithms for recommending predicates (RDF statements) about resources, based on statistical dataset analysis (i.e. on similarity and co-occurrence). Both algortihms are implemented in ActiveRDF.

Johanna Völker from AIFB Karlsruhe concludes the session with 'Learning Disjointness'. So, why disjointness is important? Of course disjointness axions allows infering new knowledge or modelling errors. But, most ontologies today don't provide disjointness axioms. Based on machine learning, an approach is presented to automated generation of disjointness axioms to be included into ontologies.

Lunchbreak -- as being in Austria I tried the traitional 'Schnitzel' (compared to German Schnitzel, Austrian Schnitzel are very flat :) )

Keynote 2 - 14.00 - 15.00
The afternoon session starts with a keynote given by Ning Zhong from Knowledgne Information System Laboratory on Ways to Develop Human-Level Web Intelligence: A Brain Informatics Perspective. Interesting thing, because I am very curious to know what 'Brain Informatics' should be... First, it's about Web Intelligence (WI). Short Definition: Web Intelligence = Artificial Intelligence + IT. Brain Informatics - on the other hand - is a new interdisciplinary field to study human information processing mechanisms systematically. It's on the intersection of cognitive science, neuro science and WI. The (H)uman brain is regarded as a (I)nformation (P)rocessing (S)ystem (HIPS). There is a gab between WI and Brain Science....
Ok....a little bit too much for me. After several (large) pictures of brains (but few content) I decided to spent the rest of the session down in the lobby. Another interesting phenomenon to mention: computer scientists gather around power outlets (at least in the afternoon, when all the batteries are getting low...)

Best paper Award Session - 15.00 - 16.00
Ok...the battery is back at 42% and the sessions continue with the best paper awards. This time, there are two best papers. The first is presented by Renee Witte from IPD Karlsruhe on 'Empowering Software Maintainers with Semantic Web Technologies'. One of the main challenges in software maintenance is to establish and maintain the semantic connections among all the different artifacts. Basically, the main contibution of the paper is on how Semantic Web technologies can deliver a unified representation (ontologies) to explore, query and reason about a multitude of software artifacts, mainly source code and other ducoments, for being deployed in security analysis, architectural evolution, and traceability recovery between source code and documents....Alas, what is still missing is a standard semantic enabled software development enviroment (e.g. as eclipse plugin...).
The other best paper is presented by Claudio Gutierrez on 'Minimal Deductive Systems for RDF'. We all use RDF for basic knowledge representation. Ofcourse - at least at first sight - RDF documents do look a little bit...lets say...complicated. But, have you ever thought about RDF containing redundancy that can be removed? That's what the authors propose by abandoning some of the RDF predicates while in the same time maintaining its semantic expressiveness. Thus, they obtain a streamlined fragment of RDFS which includes all the vocabulary that is relevant for describing data, avoiding vocabulary and semantics that theoretically corresponds to the definition of the structure of the language. Ok....but do we really want to abandon RDF and substitute it with this RDF core? As brilliant as the paper and the idea is, I guess (and so did also one of the questions to Claudio) in practice it will be only of little relevance..

Personalization I / Social Semantic Web - 16.30 - 18.00
For the 2nd afternoon session I have to switch between two of the parallel sessions. I start with the Personalization session and after the first talk I switch to Social Semantic Web.
Alistair Duke from Next Generation Web Research Group presents a talk on 'Squirrel: An Advanced Semantic Search and Browse Facility'. Squirrel is a tool for searching and browsing semantically annotated data. To achieve this, Squirrel provides combined keyword based search (providing ease of use and speed) and (powerful) semantic search in a hybrid approach. They comine several interesting technologies such as ontology management, named entity recognition on ontology generation and classification to enable features such as result consolidation, natural language generation, ontology based user profiling and device independence.

Ok...searching for the 'Saal Strassburg' was a little bit difficult (because of the missing signs). Thus, I am a little bit late for Jerome Euzenat's talk on 'Towards Semantic Social Networks'. He introduces a Semantic Social Network theory, based on a threelayered model, which involves the network between people (social network -> social layer), the network between the ontologies they use (ontology network -> ontology layer) and a network between concepts occurring in these ontologies ( -> concept layer). He proposes a similarity measure between concepts and is propagating this similarity to a distance and an alignment relation between ontologies (using these concepts). In addition, this distance relation can be used for discovering affinity in the social network (finding people that think the same way as you do...possibly).
The last talk of today is given by Nicholas John Kings on 'Knowledge Sharing on the Semantic Web'. 'The Semantic Web is Dead....', people just do tagging. With a slide like this, for sure you will get some attention (at a Semantic Web conference)...and Nicholas did :) After this eye-catching introduction, he introduces 'Squidz' a tool for automatically classifying browsed web pages against an ontology ('Spyware on Steroids...')), and allowing users to share comments made about those pages to members of a community. After all, the project is intended to test the hypothesis that information sharing is more effective when the software is aware of both the social and technical context of that information. Thus, leading to a hybrid approach combining folkosonies (free annotations of web pages) with formal ontologies.

So....just looking forward to the poster session (and the included dinner of course) and another power outlet, because the batteries are down to 8%(!).

[ok...switching to day 02 .... again trying to achieve 'live blogging'...]

Wednesday, March 28, 2007

The State of the Semantic Web


In the DERI blog I have found a rather interesting presentation from Ivan Herman, (W3C) Semantic Web Activity Leader, about the current state of the semantic web, given at the International Conference on Semantic Web & Digital Libraries in Bangalore, India.

He also gives some insight in what has been achieved concerning language standardization (RDF, OWL, SPARQL, etc.), already existing vocabularies, current implementations and available knowledge databases (ontologies).

Besides all the achievments of the recent years, Ivan also points out where the deficiencies are:

  • how to bind to different communities (e.g., the “digital library world”)

  • how (and where from) to get RDF data

  • several stil missing functionalitiessuch as rules, “light” ontologies, fuzzy reasoning, necessity to review RDF and OWL,…

  • different misconceptions, existing messaging problems

  • and of course the need for more applications, more deployment, and also acceptance


To get more working applications the porting of alreay existing knowledge resources (as already existing classification schemes, lexicons, htesauri, etc.) into the context of the semantic web is mandatory. SKOS (Simple Knowledge Organization System) is one of the W3C activities that has exactely this purpose. SKOS is built on top of RDF and serves as a formal language for representing structured well defined vocabularies.

Tuesday, March 20, 2007

Liquid Browsing - a new paradigm for accessing huge collections of data ?!


As promissed yesterday, I've installed iverse's liquidfile, a kind of substitute (or supplement) for the mac finder. (For all those mac illiterates: the finder is some kind of file system browser). Liquidfile applies the principle of liquid browsing to the filesystem of your computer. To get a short overview, how liquid browsing works, just take a look at the following short movie presentations [1] [2]. It's rather difficult to describe the way how it works with words, because its a visual way of browsing...and thus it is better explained in a visual way also. It's a 2 dimensional visualization, where e.g. the x-axis represents a timeline (as e.g. file creation date) and the y-axis represents all files in an alphabetical order. Then, every file is put on that grid and denoted with a bubble. The size of the bubble (which itself is semi-transparent) reflects the size of the file. The nice thing in general about liquid browsing is that you are able to visualize huge amount of data also on small displays. In that case the mouse pointer acts as some kind of magnifying glass.

As for the file system on your computer, this kind of visualizations has some advantages. First, you can dissolve the entire directory structure (if you want) and look at all files at once (on my computer this were more than 15.000 files). You have the possibility to combine complex (realtime) filtering with several selection mechanisms to find the files you are looking for. Simultaneously, the order in which the files are presented in general stays always the same and does not change. Thus, your visual memory will always be able to memorize the approximate position of a distinct file that you are looking for.
Of course the system has several drawbacks (at least now). My macBook is rather new (january 2007). Thus, almost all of my files have the same creation date (most of them were simply copied to the macBook on the very same day....). Therefore, the time axis will work best, if you start to work with your computer and they will be created/changed over the time.
Another drawback is the limitation of the axes to represent only filenames, filesizes, or creation/change dates. I would like to order files also according to other (also content based) criterias to get a better overview.

But, it's a nice appetizer anyway. What I would like to try out is to visualize relationships between entities (files/documents) with this technique. Just imagine bubbles of different entities and if I focus on one, all other bubbles that represent entities, which are in a certain relationship to the entity being in focus will highlight or move closer (of course in realtime...).

Wednesday, November 15, 2006

wikipedia to serve as a global ontology....



Today, I met Lars Zapf for a quick coffee enjoying the rare late afternoon november sun. We were exchanging news about ISWC, WebModay, recent projects, and stuff like that. While talking about semantic annotation, Lars pointed out that instead of using (or developing) own ontologies for annotating (and authoring) documents, you could also use a wikipedia reference to indicate the semantic concept that you are writing about. Thus, as he already wrote in a comment, e.g., you could use the link http://en.wikipedia.org/wiki/Rome to indicate that you are refering to the city of Rome, the capital of Italy.
Of course you might object that there are several language versions of wikipedia and thus, there are several (different) articles that refer to the city of Rome. To use wikipedia as a 'commonly agreed and shared conceptualization' - to fulfill at least some points of Tom Gruber's ontology definition as long as wikipedia lacks the 'formal' aspect of machine understandability - we can make use of the fact that articles in wikipedia can be identified with articles in other language versions with the help of the language indicators at the lower left side of wikipedia's user interface. To serve as a real ontology, each wikipedia article should (at least) be connected to formalized concept (maybe encoded in RDF or OWL). This concept does not necessarely have to reflect all the aspects that are reported in the natural language wikipedia article. E.g., Semantic Media Wiki is working on a wiki extension to capture simple conceptualizations (such as e.g. classes or relationships).
An application for authoring documents could easily be upgraded by offering links to related wikipedia articles. If the author enters the string 'Rome', the application could offer the related wikipedia link to Rome [or any selection of related offers] and according to the authors this link can be automatically encoded as a semantic annotation (link).
O.k., that sounds pretty simply. Are any students out there to implement it (anybody in need for credit points??)? I would highly appreciate that...

Thursday, November 09, 2006

International Semantic Web Conference 2006 (ISWC 2006), Athens (GA), USA - Day 2


Wednesday...the 2nd day of ISWC started with a keynote of Jane E. Fountain from the University of Massachussetts in Amherst about 'The Semantic Web and Networked Governance'. From her point of view, Governements have to be considered as major information processing [and knowledge creating] entities in the world, and she was trying topoint out the key challenges faced by governements in a networked world (for me the topic was not that interesting...). Also today's sessions - at least those that I have attended - were not that exciting. I liked one presentation given by Natasha Noy from Stanford on 'A Framework for Ontology Evolution in Collaborative Environments' in the 'Collaboration and Cooperation' session. She presented an extension of the protégé ontology editor for collaborative ontology development.
The most interesting session for me was the 'Web 2.0' panel in the afternoon. Amon the panelist were Prof. Jürgen Angele (Ontoprise), Dave Beckett (Yahoo!), Sir Tim Berners-Lee (W3C), Prof. Benjamin Grosof (MIT Sloane School of Management), and Tom Gruber. The panelwas discussing the role of semantic web technology for web 2.0 applications.


Jürgen Angele pointed out that the only thing that is really new about web 2.0 is ad-hoc remixability. Everything else is nothing but 'old' technology. But, as he stated, web 2.0 could be a driving force for semantic web technology.

Dave Beckett made some advertising for Yahoo! in the sense that he was pointing out that Yahoo! indeed is making use of semantic web technology (at least in their new system called Yahoo!Food) and Yahoo! is a great participation platform with more than 500 million visitors per month.

Tim Berners-Lee gave a survey on the flaws and drawbacks of web 2.0 and how semantic web technology could help. While web 2.0 is not able to provide real inter-application integration, the semantic web on the other side does not provide such cool interfaces to data. Together both in combination, they could become interesting.
All so called new aspects of web 2.0 have already been the goals of the original web (1.0), as easy creation of content, collaborative spaces, intercreativity, collective intelligence from designing together, creating relationships, reuse of information, and of course user-generated content. Web 2.0 architecture consists of client side (AJAX) interaction and server side data processing (aka the good old 'client-server'-paradigm) and mashups (one per application / each needs coding in javascript, each needs scraping/converting/...). Essentially, web 2.0 is fully centralized. So, why are skype, del.icio.us, or flickr websites instead of protocols (as foaf is)? The reuse of web 2.0 data is only limited to the hostside. Only with the help of feeds, data are able to break out from centralized sites. What will happen with all of your tags? Will they end up as simply being words or will they become real (and usefull) URIs?
With semantic web technology, web 2.0 enables multiple identities for you. You may have many URIs, enabling you to access different sorts of data, to fullfill different expectations concerning trust, accuracy, and persistence. In the end, web 2.0 and semantic web while being good seperately could be great together!

Benjamin Grosof asked, where semantic web technology could help web 2.0. He focused on backend semantic integration and mediation (augment your information via shallow inferences), collaboration and semantic search. Semantic search will enable you a morhuman centered search interface, as e.g., 'Give me all recipes of cake....but I don't like any fruits' and 'I want a good recommendation from a well reputed web site'. He sees semantic web technology piggyback on web 2.0 interactions ('web 2.0 = search for terrestrial intelligence in the crowd' :) The semantic web should exploit web 2.0 to obtain knowledge.

Tom Gruber was asking 'Where is the mojo in Web 2.0?'. He characterized web 2.0 as being a fundamentally democratic architecture, driven by social and entertainment payoffs (universal appeal...), while the web 1.0 business model actually keeps working ('attention economy'). He was discussing the way from today's 'collected intelligence' to real 'collective intelligence'. He concluded 'don't ask what the web knows....ask what the world knows!' and 'don't make the web smart...make the world smart'.

Wednesday, November 08, 2006

International Semantic Web Conference 2006 (ISWC 2006), Athens (GA), USA - Day 1


Tuesday morning 9 a.m. ... the ISWC 2006 starts with the keynote of Tom Gruber (godfather of computer science based definition of the term 'ontology') on 'Where the Social Web Meets the Semantic Web'. He focused on 'Collective Intelligence' as being the reason that companies as google or amazon did survive the first Dot-com bubble, because they where making use of their users' collective knowledge. Google uses other people's intelligence by computing a page rank out of the users' links to other webpages. Amazon uses the people's choices for their recommentation system, and ebay uses the people's reputation. Interesting thing about that is that the notion of 'Collective Intelligence' (aka 'Social Web', aka 'Web 2.0') - was already addressed by Douglas Engelbart in the late 60's. Engelbart did not only invent the mouse, the window-based user interface, and many other important things that are part of today's computing environment, his driving force - as Gruber said - was 'Collective Intelligence'....to cope with the set of growing problems that humanity is facing today. Thus - as I have also stated in another post - also the semantic web depends on collaboration and participation of the users and therfore, on 'Collective Intelligence' to become a success.

BTW, I prefer using the term 'Social Web' instead of 'Web 2.0'. From my point of view 'Social Web' hits exactly the point and does not suggest any new and exciting technology (but only the fact that people are using existing web technology in a collaborative way to interact with each other).

After the keynote I visited the 'Knowledge Representation' session with an interesting talk of Sören Auer on OntoWiki (a semantic wiki system .. interesting, because one of my students is alsoimplementing a semantic wiki). In the afternoon sessions I esp. liked the talks about representation and visualization (esp. the talk of Eyal Oren on 'Extending faceted navigation for RDF data', where he presented a nice server application that is able to visualize arbitrary RDF-data). In the evening, a dinner buffet (including cuban music) was combined with the poster sessions and the 'Semantic Web Challenge' exhibition, where I found the possibility for a cooperation with Siegfried Handschuh from DERI (on semantic authoring and annotation....).

Oh...I already forgot to mention that there is also a flickr group with ISWC photographs...