Showing posts with label 2007. Show all posts
Showing posts with label 2007. Show all posts

Monday, June 04, 2007

ESWC 2007 - European Semantic Web Conference, Innsbruck (Day 01)



Today, the 4th European Semantic Web Conference starts in Innsbruck (Austria). I arrived already yesterday in the evening by train. By chance, I already met 2 other participants (Pascal Hitzler and Andreas Hotho ... -> bibsonomy) while waiting at Munich train station. Thus, the last part of the train ride was quite entertaining :)

Keynote 1 - 9.00 -10.00
Day 1 of ESWC 2007 is starting with a keynote of Stefano Ceri on 'Design Abstractions for Innovative Web Applications: the case of the SOA augmented with Semantics'. Sorry to say, but -- at least for me -- the talk was not really interesting. Since I am not so much interested in Semantic Web Services, I decided to spent the rest of the morning sessions (after the coffee break) in the Ontology Learning, Inference and Mapping session. In parallel there are session on Semantic Web Services and Semantic Web Use Cases (unfortunately in areas also not very interesting for me, i.e. mechatronics and e-Governement). But, I'm looking forward to the Ontology Engineering session....

Ontology Engineering I - 10.30 - 12.30
The Ontology Engineering session starts with Martin Hepp on 'GenTax: A Generic Methodology for Deriving OWL and RDF-S Ontologies from Hierarchical Classifications, Thesauri, and Inconsistent Taxonomies', where he is referring to a novel methodology for automatically deriving consistent RDF-S and OWL ontologies from informal hierarchies. He demonstrated the usefullness of the approach by utilizing it for transforming the two e-business categorization standards eCl@ss and UNSPSC into ontologies. The approach is based on drawing a random sample from an input classification, letting a human user confirm/reject/adapt the sample, and using this confirmed sample as a basis for further automated classification. The tool will be available here soon.

The second talk is presented by Maciej Janik on 'SPARQLeR: Extended Sparql for Semantic Association Discovery'. Semantic Associations (i.e. complex semantic relationships among entities) in knowledge bases have to be discovered, which comes to the process of finding paths of possibly unknown length that connect the given entities and have a specific semantics. SPARQL for querying RDF databses is extended for the discovery of such paths including the possibility of formulating regular expressions over properties for specifying the required semantics of the queried paths.

Eyal Oren from DERI Galway is speaking about 'Algorithms for Predicate Suggestions using Similarity and Co-Occurrence'. Only through shared vocabularies can meaning be established. Tagging Systems solve this problem by tag suggestions to achieve coherrent vocabulary. The same principle holds for creating distributed RDF databases. Therefore, two domain-independent algorithms for recommending predicates (RDF statements) about resources, based on statistical dataset analysis (i.e. on similarity and co-occurrence). Both algortihms are implemented in ActiveRDF.

Johanna Völker from AIFB Karlsruhe concludes the session with 'Learning Disjointness'. So, why disjointness is important? Of course disjointness axions allows infering new knowledge or modelling errors. But, most ontologies today don't provide disjointness axioms. Based on machine learning, an approach is presented to automated generation of disjointness axioms to be included into ontologies.

Lunchbreak -- as being in Austria I tried the traitional 'Schnitzel' (compared to German Schnitzel, Austrian Schnitzel are very flat :) )

Keynote 2 - 14.00 - 15.00
The afternoon session starts with a keynote given by Ning Zhong from Knowledgne Information System Laboratory on Ways to Develop Human-Level Web Intelligence: A Brain Informatics Perspective. Interesting thing, because I am very curious to know what 'Brain Informatics' should be... First, it's about Web Intelligence (WI). Short Definition: Web Intelligence = Artificial Intelligence + IT. Brain Informatics - on the other hand - is a new interdisciplinary field to study human information processing mechanisms systematically. It's on the intersection of cognitive science, neuro science and WI. The (H)uman brain is regarded as a (I)nformation (P)rocessing (S)ystem (HIPS). There is a gab between WI and Brain Science....
Ok....a little bit too much for me. After several (large) pictures of brains (but few content) I decided to spent the rest of the session down in the lobby. Another interesting phenomenon to mention: computer scientists gather around power outlets (at least in the afternoon, when all the batteries are getting low...)

Best paper Award Session - 15.00 - 16.00
Ok...the battery is back at 42% and the sessions continue with the best paper awards. This time, there are two best papers. The first is presented by Renee Witte from IPD Karlsruhe on 'Empowering Software Maintainers with Semantic Web Technologies'. One of the main challenges in software maintenance is to establish and maintain the semantic connections among all the different artifacts. Basically, the main contibution of the paper is on how Semantic Web technologies can deliver a unified representation (ontologies) to explore, query and reason about a multitude of software artifacts, mainly source code and other ducoments, for being deployed in security analysis, architectural evolution, and traceability recovery between source code and documents....Alas, what is still missing is a standard semantic enabled software development enviroment (e.g. as eclipse plugin...).
The other best paper is presented by Claudio Gutierrez on 'Minimal Deductive Systems for RDF'. We all use RDF for basic knowledge representation. Ofcourse - at least at first sight - RDF documents do look a little bit...lets say...complicated. But, have you ever thought about RDF containing redundancy that can be removed? That's what the authors propose by abandoning some of the RDF predicates while in the same time maintaining its semantic expressiveness. Thus, they obtain a streamlined fragment of RDFS which includes all the vocabulary that is relevant for describing data, avoiding vocabulary and semantics that theoretically corresponds to the definition of the structure of the language. Ok....but do we really want to abandon RDF and substitute it with this RDF core? As brilliant as the paper and the idea is, I guess (and so did also one of the questions to Claudio) in practice it will be only of little relevance..

Personalization I / Social Semantic Web - 16.30 - 18.00
For the 2nd afternoon session I have to switch between two of the parallel sessions. I start with the Personalization session and after the first talk I switch to Social Semantic Web.
Alistair Duke from Next Generation Web Research Group presents a talk on 'Squirrel: An Advanced Semantic Search and Browse Facility'. Squirrel is a tool for searching and browsing semantically annotated data. To achieve this, Squirrel provides combined keyword based search (providing ease of use and speed) and (powerful) semantic search in a hybrid approach. They comine several interesting technologies such as ontology management, named entity recognition on ontology generation and classification to enable features such as result consolidation, natural language generation, ontology based user profiling and device independence.

Ok...searching for the 'Saal Strassburg' was a little bit difficult (because of the missing signs). Thus, I am a little bit late for Jerome Euzenat's talk on 'Towards Semantic Social Networks'. He introduces a Semantic Social Network theory, based on a threelayered model, which involves the network between people (social network -> social layer), the network between the ontologies they use (ontology network -> ontology layer) and a network between concepts occurring in these ontologies ( -> concept layer). He proposes a similarity measure between concepts and is propagating this similarity to a distance and an alignment relation between ontologies (using these concepts). In addition, this distance relation can be used for discovering affinity in the social network (finding people that think the same way as you do...possibly).
The last talk of today is given by Nicholas John Kings on 'Knowledge Sharing on the Semantic Web'. 'The Semantic Web is Dead....', people just do tagging. With a slide like this, for sure you will get some attention (at a Semantic Web conference)...and Nicholas did :) After this eye-catching introduction, he introduces 'Squidz' a tool for automatically classifying browsed web pages against an ontology ('Spyware on Steroids...')), and allowing users to share comments made about those pages to members of a community. After all, the project is intended to test the hypothesis that information sharing is more effective when the software is aware of both the social and technical context of that information. Thus, leading to a hybrid approach combining folkosonies (free annotations of web pages) with formal ontologies.

So....just looking forward to the poster session (and the included dinner of course) and another power outlet, because the batteries are down to 8%(!).

[ok...switching to day 02 .... again trying to achieve 'live blogging'...]

Tuesday, April 17, 2007

Summer semester starts at FSU Jena...

Yesterday, the summer semester started and hey....it's really summer! I can't remember any summer semester start with such high temperatures. This semester, I have to give 2 lectures: 'Einführung in die Informatik II' (introduction to computer science II), which takes place mondays and fridays, and 'Informatik der digitalen Medien' (computer science and digital media), also on mondays. The later wil be recorded and will soon be available also for offline participation :)
Intersting thing to mention, this semester I try to maintain a blog for each of the lectures for providing additional information, reading materials, bibliographic notes, etc. You will find the two new blogs at
Of course I will comment on my experiences about managing lectures and courses with a blog.

Wednesday, January 24, 2007

SOFSEM - Day 4

Now we have snow....finally :) ...even a lot of it. It was snowing all day long, roads in Czech Republic and also in southern Germany were closed. Also Prague Airport was closed until the afternoon. But, I guess as far as I remember that are the more typical weather conditions for SOFSEM.
Anyway, the day started with a keynote of Tom Henziger about 'Games, Time, and Probability: Graph Models for System Design and Analysis'. He addressed three major sources of system complexity: concurrency, real time, and uncertainty. Concurrency can be modelled as a multi-player game representing a reactive system with potential collaborators and adversaries. Real time requires the system to combine discrete state changes as well as continous state evolution, while state changes - for uncertainty - also have to be modelled in a probabilistic way.
Unfortunately some of the presenters of the following contributed papers did not show up. Thus, the conference program was subject to several changes. In the afternoon the posters of the student research forum each had a short 5 minute presentation, followed by a poster exhibition and a lot of discussions. In the end, the participants should give a vote for the best poster presentation. My choice - which of course is completely subjective - was the poster of of Henning Fernau and Daniel Raible on 'Alliances in Graphs: a Complexity-Theoretic Study'.
In the late evening I was trying to look for my car, which was buried under the snow at the parking lot. Due to the wind the snow around the parking lot (and my car) was piled up almost half a meter...which made me think about the road conditions and the plan of driving home the next day....

Tuesday, January 23, 2007

SOFSEM 2007 - Day 3

Today started with a keynote given by Ricardo Baeza-Yates from Yahoo! Research on 'Mining Web Queries'. In particular he showed how to identify categories of user queries and how to use this information to create an appropriate ranking of the search results. Besides the already identified 'coarse' categories, such as, e.g., queries being 'informational', 'navigational', or 'transactional' (which means that the user wants to have (a) information about a specified topic, (b) a starting point for further research, or (c) a homepage related to the resource for transactional purposes (e.g. shopping)...), he addressed several graphs that can be compiled out of the search engine logfile, as e. g., URL cover graph, URL link graph, session graph...These graphs can be used for identifying polysemic expressions, similar or related queries, clusterings of queries, or even a (pseudo)taxonomy of queries.
Besides web query mining, he mentioned some interesting numbers concerning Yahoo, as e.g. that Yahoo administrates about 20 PetaBytes of Data with more than 10 TeraBytes of data traffic per day. But, on the other hand, he gave an estimation of the actual world knowledge and related it to the ammount of data managed by Yahoo today: given that a person creates about 10 pages of data concerning a distinct event, and if we estimate the number of events of about 5000 in a lifetime, and if we multiply that number by the world's population....we will end up with about 0,0057% of the 'world knowledge' currently being represented in Yahoo...

Monday, January 22, 2007

SOFSEM 2007 - Day 2


The second day of SOFSEM started with a keynote of Bertrand Meyer (maybe you remember Eiffel...) from ETH Zürich on 'Automatic Testing of Object-Oriented Software'. To enable automated testing, he referred the concept of 'contracts' being directly embedded in the classes of the Eiffel programming language. With a contract you are able to specify the software's expected behaviour (preconditions, postconditions, and invariants). which can be monitored during execution. In automated software testing, contracts may serve as test oracles that decide, whether a test case has passed or failed. He presented 'Auto Test' unit testing framework, which is using Eiffel contracts as test oracles. Auto Test is able to exercise all classes by generating objects and routine arguments. Also manual testing can be embedded as well as regression testing for failed test cases, which is implemented in a 'minimized' form by retaining only the relevant instructions.

For the rest of the second day contributed (refereed) paper presentations are scheduled. I will have to chair the first session of the 'emerging web technologies' track, which will be on XML technology. If there (or in any other session I attend) will be anything of interest, you will read it right here ... :)
So...Joe Tekli from the Université de Bourgogne presented a 'Hybrid Approach on XML-Similarity', which combined structural similarity of XML-Documents with 'semantic' arguments, i.e. tag names of different XML-documents are compared with the help of WordNet to compute some similarity measure. Quite an interesting application that can be build on, esp. regarding the semantic similarity aspect. But nevertheless, maybe we can use it for our MPEG-7 based video search system (OSOTIS).

Sunday, January 21, 2007

SOFSEM 2007 - Day 1


This year, after about 7 or 8 years, I am attending again the SOFSEM conference on 'Current Trends in Theory and Practice of Computer Science' (for the 2nd time). Maybe SOFSEM is not the most important of all the computer science conferences around, but it is rather original and has quite some history (i.e. it's tradition dates back more than 30 years...). SOFSEM means SOFtware SEMinar, and this already gives some hint about its originality. Starting from a winter lecture with only limited international attendance it has developed to an interesting mixture of lectures (given by invited speakers of significant reputation), presentations of reviewed research papers, and student paper presentations. By tradition, it's location always switches between somewhere in Slowakia and the Czech Republique and always in winter. Unfortunately, this year winter did not really show up and thus, we are sitting here in Harrachow (a well known winter resort) without any snow. On the other hand, nice thing about this situation is that travelling this year has become much easier (because there is no snow even in the mountain areas).
This year, I am co-chairing the track 'emerging web technologies' as being one of the four SOFSEM tracks. By tradition, there is always a track 'foundations of computer science' besides of three changable tracks concering breaking topics of current interest , i.e. (in this year) 'multi-agent systems', 'emerging web technologies', and 'dependable software and systems'.
The first day on SOFSEM, after the opening note given by Jan van Leeuwen, in which he referred to the long tradition of SOFSEM and to Czech computer science history, starts with a full day of invited lectures covering all four topics.

  • Manfred Broy from TU Munich started with a presentation on 'Interaction and Realizability'. In interactive computation - in difference to sequential, atomic computation - input as well as output is not provided as a whole, but step by step while the computation continues. He pointed out that interactive behaviour can be modeled with Moore machines and introduced the term of 'realizability', which is a fundamental issue when asking whether a behaviour corresponds to a computation. 'Realizable functions' are defined as being abstractions of state machines (in a similar way as partial functions are abstractions of Turing machines) and can be used to extend the idea of computability to interactive computations.

  • Andrew Goldberg followed with a talk on 'Point-to-Point Shortest Path Algorithms with Preprocessing'. To run on even small devices while at the same time covering graphs with tens of millions of nodes (as, e.g., in roadmaps for navigation devices), off course efficient algorithms are required. The traditional way is to search a ball around the starting point (as e.g. in Dijkstra's algorithm) that can be speed up by biasing the search towards to intendet target point (as e.g. in A* search, if additional information is available that provides a lower-bound on the distance to the target) or by pruning the search graph (as e.g. in ALT algorithms that precompute distances to preselected landmarks, or using 'reaches').

  • Jerome Lang from IRIT (France) continued the afternoon session with a survey on 'Computational Issues in Group Decision Making', which combines 'social choice' (from economics) and AI (applications) into 'computational social choice' theory. In this new and very active discipline concepts as e.g. voting procedures, coalition formation, and fair division (from social choice), which is also important for multi-agent systems, are examined under the consideration of complexity analyses and algorithm design.

  • I realized that I will be the chairman of today's last session. Thus, the summary of Remco Veltkamp's (University of Utrecht, The Netherlands) talk on 'Multimedia Retrieval Algorithms' will come with a little delay....
    The presentation started with citing Marshal McLuhan's famous quote 'The medium is the message' smartly being connected to the basic definitions of multimedia retrieval. Difficult thing in multimedia retrieval is the proper understanding of the mechanisms of human perception and in connection to that the question of how to take care of it's peculiarity in information retrieval. E. g., the human visual system is famous for 'generic interpretations', i.e. sometimes we see things that are not really there, as already has been described by Wertheimer's Gestalttheorie back in 1923. Interesting fact, that some of these visual illusions do also exist for audio perception. For multimedia retrieval metrics have to be defined for computing similarities (as well as differences of multimedia objects) in an efficient way, while the algorithms dealing with multimedia retrieval have to be carefully designed according to the type of problem that is addressed (e.g., computing problem, optimization problem, decision problem, etc.). The presentation closed with a short demonstration of the music search engine Muugle that realizes the concept of 'query-by-humming'.