Showing posts with label information retrieval. Show all posts
Showing posts with label information retrieval. Show all posts

Monday, October 26, 2009

Open PhD Positions in Semantic Multimedia Retrieval Project

OPEN Ph.D. POSITIONS at Hasso-Plattner-Institute (HPI), Potsdam (Germany) starting on the fourth quarter of 2009

Hasso-Plattner-Institute (HPI) is a privately financed institute affiliated with the University of Potsdam, Germany. The Institute's founder and benefactor Professor Hasso Plattner, who is also co-founder and chairman of the supervisory board of SAP AG, has created an opportunity for students to experience a unique education in IT systems engineering in a professional research environment with a strong practice orientation.
(for more information on HPI, c.f. http://www.hpi.uni-potsdam.de/ )

Project Description:
MEDIAGLOBE is part of the THESEUS research program initiated by the German Federal Ministry of Economy and Technology (BMWi), with the goal of developing a new Internet-based infrastructure in order to better use and utilize the knowledge available on the Internet. The focus of the research program is on semantic technologies, which determine contents (words, images, sounds, and videos) not through conventional methods (e.g., combinations of letters) but which are able to recognize and place the meaning of a content in its proper context. MEDIAGLOBE deals with digitalization, analysis, and semantic retrieval of historical, documentary audiovisual content. (for more information on MEDIAGLOBE, c.f. http://theseus-programm.de/theseus-mittelstand-2009/ )

The ideal candidate holds a MS degree in Computer Science or related field and is able to consider both theoretical and practical/implementation aspects in her/his work. Fluent english communication and programming skills are fundamental requirements. Since we are working on a multimedia repository with resources in German language, German language skills are welcome! Preferably the candidate has a background in one of the following
fields:
• semantic web technologies
• knowledge representations and ontology engineering
• audiovisual retrieval and analysis
• semantic search
• innovative web development
• user interface design for audiovisual content

The position starts as soon as possible and is full-time (40h/week) for the duration of the project until Oct 2011. Review of applications will begin immediately and will continue until the position is filled. The successful candidate will tightly work with international partners and has the possibility to pursue PhD work within the scope of the project.

How to apply:
Excellent candidates are invited to apply with:
• Curriculum vitae and copies of degree certificates/transcripts,
• Writing samples/copies of relevant scientific papers (e.g. thesis, etc.),
• Letters of recommendation.

Please send your application in PDF format indicating in the subject 'Application for PhD position‘ via email or via traditional mail to the following contact.

Contact and application:
Harald Sack
Hasso-Plattner-Institut für Softwaresystemtechnik GmbH
Universität Potsdam
Prof.-Dr.-Helmert-Str. 2-3
D-14482 Potsdam, Germany
phone: +49 (0)331-5509-527
fax:
+49 (0)331-5509-325
email:
harald.sack@hpi.uni-potsdam.de
web:
http://www.hpi.uni-potsdam.de/meinel/persons/sack.html

Friday, June 27, 2008

Adaptive Multimedia Retrieval 2008 in Berlin, June 26-27, 2008 - Day 02

After the dinner cruise along the river Spree, the second day of Adaptive Multimedia Retrieval 2008 again starts with an interesting invited talk on the European answer to Google search engine technology - THESEUS.

Karsten Müller from Fraunhofer Heinrich-Hertz-Institute is presenting on "THESEUS Project - Applications and Core Technologies for the Semantic Web". First, Karsten makes clear, that THESEUS doesn't want to be Google ;-) THESEUS is a research program for a new internetbased knowledge infrastructure....which from my point of view means nothing else but "the semantic web"....
One part of the THESEUS project is ALEXANDRIA, the virtual library, being lead by Yahoo! with the objective of semantic processing of different forms of content to enable faster access to relevant content, which again means an increase in information quality. Concepts such as an automated tagging framework (including language error correction, synonym & tag merging, and topic focussing, identification of semantic relations), innovative navigation (by presenting thematically related contents) and interaction concepts are involved.
Another part is ORDO, which deals with "Organizing your digital life" with the goal to unify various data formats, multilingual information, structured and unstructured data on the web to enable homogeneous information sources.Problems such as separating important from unimportant, ordering information instead of searching, priorization, identification and visualization of interrelations are addressed.
TEXO is another part with the objective of "Realizing the internet of services" (being lead by SAP Research), offering personalized customized services, community involvement to improve services, as well as a smooth & seamless (userfriendly) adaption and integration of services.
PROCESSUS deals with the "Optimization of business processes" aiming for the objective of anytime providing the user with theright information at any stage of the business process.
MEDICO is another subproject dealing with "Towards Scalable Semantic Image Search in Medicine" and being lead by Siemens.
CONTENTUS, as being the last Use case "Content access and generation from cultural institutions is lead by the Deutsche Nationalbibliothek. Being part of CONTENTUS are tasks such as Digitizing books as well as audiovisual material (including the German Music Archive in Berlin) protecting the cultural heritage. The goal is the semantically interlinked collection of content to achieve a next generation multimedia library.
.....impressive and ambitious project!

The upcoming section this morning is on "Image Tagging" and Marius Renn (at least I hope so) from TU Kaiserslautern is givig a presentation on "Automatic Image Tagging using Community-Driven Online Image Database". Automatic image tagging requires a lot of training data....and flickr is delivering tons of tags per day...but are these flickr data really good candidates for learning? So, in the end, unfiltered community image sets directly do not provide satisfying results. Alas, these databases at least allow large scale image aggregation...
The next talk in this session is given by Christian Hentschel from Fraunhofer HHI Berlin about "Automatic Image Annotation Refinement using Object Co-Occurences". Again, flickr is the target image set with its huge collection of more than 2 billion images, growing by 3 million photos every day. Objects always appear and are perceived in a semantic context.

The following session is on "Symbolic Music Retrieval" and starts with Rainer Typke from Austrian Research Institute for Artificial Intelligence (ÖFAI), but I had to skip this talk. Anyway, the samples of the reduced MIDI files were quite interesting (although I'm not a fan of the Scorpions!). OK, I had to ask afterwards about the usefulness and application of his approach. In music retrieval it can be used to reduce the index size down to 30% of the original index. Also QBE-processing will become much easier while on the other hand you might connect this MIDI-collection to real music files.
The last talk of the morning session is given by Giancarlo Vercellesi from University of Milan on "Automatic synchronization between audio and partial music score presentation". He presents the ParSi architecture, which perfoms an alignment of PCM signal and partial MIDI scores.

The afternoon session is simply entitled with "Systems". Fernando Lopéz from Madrid is giving a presentation on "Towards a fully MPEG-21 compliant adaption engine: complementary description tools and architectural models". Within the MPEG-21 framework several aspects of metadata-driven adaption is not clearly covered. He introduces CAIN, a tool for adapting Digital Items e.g. to different output devices.
The session continues with a presentation on "Mobile museum guide based on fast SIFT recognition" with the objective to identify paintings in galleries simply with the help of mobile pattern recognition without any extra installation on site. The SIFT (Scale Invariant Feature Transform) method is a rather cool algorithm for detecting local features within images that are used to map photographs taken with your PDA or mobile phone in the image gallery with reference pictures from a given database. And actually the live demo did work :)
I guess, we will also use the SIFT-algorithm in yovisto for synchronization of ppt/pdf-slides with the lecture video.

For the last session - "Structuring of Image Collections" - only one speaker showed up. Marc Gelgon is presenting on "Geo-temporal structuring of a personal image database with two-level variational Bayes mixture estimation".

[to be continued...]

Thursday, June 26, 2008

Adaptive Multimedia Retrieval 2008 in Berlin, June 26-27, 2008

The next two days, we are attending the Berlin Adaptve Multimedia Retrieval 2008 Workshop at the Heinrich Hertz Institute being located in downtown Berlin. So, it's pretty close to home and the only travelling involved was by S-Bahn :)

The first speaker is Francois Pachett from Sony CSL giving a keynote entitled "What are our audio features worth?"
The fundamental questions are "What makes objects what they are?", ""What are the features of subjectivity?", "How do we perceive objects and how can we transfer this to a machine?" Pachet's research is concerned with the classification of musical objects based on the so called polyphonic timbre that describes the sum of all features of a music object. Interesting thing is the identification of hubs, i.e. songs that are pretty close to every other song. Hubs in general seem to be mere artefacts of static models.
Interesting fact ist that there are companies now, predicting if your song is going to be a hit. Their judgement also relies on feature analysis and they even give recommendations how your song can be improvent to become a hit. Of course you have to pay for that service...but does it really work??

After the coffee break, there's a session on User-Adaptive Music Retrieval. The first talak is presented by Kay Wolter from Fraunhofer IDMT Ilmenau on "Adaptive User-Modelling for Content-Based Music Retrieval". They are adapting a content-based music retrieval system (CBMR) according to user preferences that are determined by acceptances and rejections of recommended songs by the user, which is furthermore used to improve the quality of music recommendations....Reminds me somehow to Pandora or last.fm...
The second talk is presented by Sebastian Stober from Otto-von-Guericke-Universität Magdeburg on "Towards User-Adaptive Structuring and Organization of Music Collections". So, wouldn't it be nice to structure your music collection automatically...but not in the way the software tells you, but the way you like it? The presented system is based on an general adaption approach using self-organizing maps that can be adapted by user interaction.

The first afternoon session is on "User-adaptive Web Retrieval" and starts with a presentation of Florian König from Johannes-Kepler-Universität Linz on "Using thematic ontologies for user- and group-based adaptive personalization in web searching". He introduces Prospector, which is a generic meta-search layer for Google, not constrained only to web search, based on re-ranking of search results and deploying user modells based on Open Directory Project (ODP) taxonomies. As far as I have understood, the applcation is based on the carrot2 framework for open source search engine result clustering.
Next, David Zellhöfer from BTU Cottbus presents on "A Poset Based Approach for Condition Weighting". Similarity search can be determined according to different conditions w.r.t. the search query. Esp. different people have different expectations if it comes to similarity. So, condition weights have to be determined by psychological experiments.

The second afternoon session is about "Music Tracking and Tumbnailing" and starts with a presentation of Tim Pohle from Johannes-Kepler-Universität Linz on "An Approach to Automatically Tracking Music Preference on Mobile Players". Ok, so the basic problem is, someday you will get bored by the music selection on your ipod. Therefore, the goal is to remove songs that you don't like anymore and replace them with new songs that you probably will like. How do you achieve this? Well, with according user feedback, i.e. by tracking the user's decision on choosing or skipping tracks. Tracks that have recently been skipped often will be dropped and replaced by tracks that are similar (according to some feature analyses) to the remaning tracks.
Next, Björn Schuller from Technische Universität Münschen is presenting on "One Day in Half an Hour: Music Thumbnailing Incorporating Harmony- and Rythm Structure". Music thumbnailing is some really cool feature, Just imagine, your sitting in your car and you are looking for another track to hear, but your player always starts songs at the beginning and they have long and boring intros. Therefore, getting to the most interesting (or significant) part of the song immediately would really be something...

The sessions close with an invited talk given by Stefan Weinzierl and Sascha Spors on "The Future of Audio Reproduction. Technology - Formats - Applications". Promissing title, let's see.... We start with a brief history of audio recording and reproduction technology starting from the very first phonograph to modern multichannel spatial surround sound systems. So, the future seems to be real sound field synthesis (wavefield synthesis, WFS) instead of relying on psycho-acustic effects as in today's stereo. Here, an array of loudspeakers reproduces exactly the wave front of the original sound source. For transmitting signals like this, no single channels are recorded anymore, but the original sound signal (without spatial characteristics of the room where it has been recorded, because this would interfere with the characteristics of the room, where it is reproduced) including movement and position of the sound source. Besides existing VRML and MPEG-4 Audio BIFS that focus more on visual scene description than on audio scene descriptions, there is the proposal of a new modeling language for high resolution spatial sound events called ASDF (Audio Scene Description Format).

[...to be continued in Adaptive Multimedia Retrieval 2008 in Berlin, June 26-27, 2008 - Day 02]

Saturday, March 24, 2007

A short note on Semantic Search Engines


Lars featured an article on Hakia, a 'semantic search engine' in his blog today.
Contrarywise to Google, Hakia uses natural language processing (NLP) to 'understand' search queries given in natural language (and not as plain keywords). Ok, this is now new.
Just remember AskJeeves, a.k.a. Ask.com. But, in difference to that search engine, Hakia claims to perform a 'semantic search' (and not a keyword based search). If you read a little bit further in their technical description at Hakia labs you will find thet they are using a parser called 'OntoSem', which as they claim is able to perform a 'deep semantic analysis' of sentences. This parser is used for query string analysis and also for the generation of the search index. But, their search index is different to that we are used from 'traditional' search engines.
Just remember, traditional keyword based search engines extract so called 'descriptors' from web pages that are used to describe the content of those pages. These descriptors are managed within an inverted index. Thus, by accessing the inverted index with a descriptor (=keyword in query string), a list of web pages will be returned together with some weights indicating the relevance of the descriptor for the page.
So, how does the index work in Hakia? They clain, that their parser is analyzing all web pages to be indexed sentence by sentence. In effect, all possible questions that can be posed for each sentence are generated, forming what they call the 'QDEX data'. All possible questions for all sentences in all pages have to be stored in an index-like data structure (simply for fast and efficient access). If now a query string contains a question, this question is mapped against the index, resulting in a large number of 'relevant' (questions, sentences and in the end...) web pages. Now they apply a 'smart' algorithm called 'Semantic Rank' which orders the resulting list of documents according to their relevance wrt. the question given in the query string. More details about the technique is not published (at least not to my knowledge).

Thus, the only way to find out about the quality of their approach is to try out their search engine. A nice feature is that they have included small example applications where you may try out OntoSem or QDEX interactively by yourself. I have only tried OntoSem (because for QDEX you have to sign up a request) and the result was not really different from any standard NLP-Parser (and thus not really convincing. I will sign up and give QDEX a try..and of course I will post the result).

If I compare their approach to Google's, the problem is that Google is also claiming to incorporate semantic technology altogether. So, again I can only compare the results of my queries. I have tried several queries and have come to the following results:

  • in general the results from hakia do look very promising!

  • in comparison to Google results, they are not really that different this may come because of my 'queries' and thus, my results are probably not really objective).
    Let me give you an example: I was asking both (Google and Hakia) the question 'When did the Semantic Web start?'. Ok, it's not so easy to understand at all. First, it didn't really start (as an being implementation), but the topic itself startet several years ago...and so I was curious about the results:

Ok, as stated before, this is not representative at all, but I will keep an eye on Hakia anyway. As soon as I have more representative experiences (and hopefully some comments or other interesting comparisons), I will write about.

One thing in the end. After all, what I have read about Hakia and on their web site, they seem to have a completely different notion of 'Semantic Search'. What they do is to apply (enhanced) information retrieval techniques based on natural language processing. They do NOT evaluate any additional given semantic annotation (RDF, OWL, etc...), which might be addid into the web pages or connected to them. They claim (same as Peter Norwig) that the average user is much too stupid to supply semantic annotation (because for doing that he needs to be a linguist as they say).
Of course, if you are following the PLAIN trail of the W3C and encode everything by hand (HTML, XML, RDF, OWL, SWRL, etc...) then you have to be some real expert. But today, you (at least most of you) don't encode HTML by your own. Remember, there are lots of real nice WYSIWYG editors for doing that AND (even more) there are several simplifications (just think of writing your blog). In the same way there will be smart user interfaces and editors providing help in generating semantic annotations for your texts. First and most simple step is providing labels and keywords (just 'tags') for your blog posts. Most people do that...just because (1) they want to put some order into the set of their postings and (2) they want their posting to be found.

Tuesday, January 23, 2007

SOFSEM 2007 - Day 3

Today started with a keynote given by Ricardo Baeza-Yates from Yahoo! Research on 'Mining Web Queries'. In particular he showed how to identify categories of user queries and how to use this information to create an appropriate ranking of the search results. Besides the already identified 'coarse' categories, such as, e.g., queries being 'informational', 'navigational', or 'transactional' (which means that the user wants to have (a) information about a specified topic, (b) a starting point for further research, or (c) a homepage related to the resource for transactional purposes (e.g. shopping)...), he addressed several graphs that can be compiled out of the search engine logfile, as e. g., URL cover graph, URL link graph, session graph...These graphs can be used for identifying polysemic expressions, similar or related queries, clusterings of queries, or even a (pseudo)taxonomy of queries.
Besides web query mining, he mentioned some interesting numbers concerning Yahoo, as e.g. that Yahoo administrates about 20 PetaBytes of Data with more than 10 TeraBytes of data traffic per day. But, on the other hand, he gave an estimation of the actual world knowledge and related it to the ammount of data managed by Yahoo today: given that a person creates about 10 pages of data concerning a distinct event, and if we estimate the number of events of about 5000 in a lifetime, and if we multiply that number by the world's population....we will end up with about 0,0057% of the 'world knowledge' currently being represented in Yahoo...

Saturday, January 13, 2007

...against all odds


On wednesday I attended a talk given by Michael Strube from EML Research on "World Knowledge induced from Wikipedia - A New Prospect of Knowledge-Based NLP ". He was showing how the (meanwhile famous) collaborative encyclopedia can be used for information retrieval purposes in a way similar to (more traditional) online dictionaries as e.g. WordNet and - though being not well structured - provides results of almost equal quality.
First thing was that for their work, Strube and his colleague regarded each Wikipedia page as being the representation of a concept (we already had some arguments about that as you might remember...). Next, they developed some metric for similarity of concepts w.r.t. to the concept hierarchy (where the wikipedia defined 'concepts' come into play). Since 2004, wikipedia features a user defined concept hierarchy. This hierarchy of concepts also can be regarded as being a folksonomy, simply because this is not a knowledge representation carefully designed by some designated domain expert, but by the wikipedia comunity in a collaborative way. Unfortunately, the wikipedia concept hierarchy suffers exactly from that fact. From my pont of view it seems problematic to compare the proposed similarity measure (based on wikipedia concept hierarchy) with other similarity measures (based on commonly shared expert ontologies). O.k., you might argue that indeed the wikipedia concept hierarchy IS commonly shared, because it has been developed by the wikipedia community...but is the knowledge represented in wikipedia really 'common'? Just remember the diversity and manifold of Star Wars characters or Star Trek episodes in wikipedia compared with, as e.g., the history of smaller Eropean countries. As for all ontologies always the view and the knowledge of the ontology designer has to be considered. The wikipedia concept hierarchy - although partly being really appropriate - reminds me somehow to this famous literary chinese dictonary entry defining the term 'animal' which is quoted by Jorge Luis Borges. Another problem lies in the fact that the different language versions of wikipedia have developed different concept hierarchies (sic!).

In the end, I was asking how this proposed information retrieval based on wikipedia could be improved by considering a 'Semantic Wikipedia', as e.g., the Semantic MediaWiki (given that those semantic wikipedias would contain sufficient data). Instead of answering my question, Michael Strube cited Peter Norwig's argument against the Semantic Web from last years AAI2006. Just to sum up: the semantic web will not become reality because of the inability of its users to provide correct semantic annotations. But hey...this guy (Strube) was talking about wikipedia. Doesn't this argument raise any associations? Just remember the time 5 or 10 years ago. Nobody (well almost nobody) would have believed that it will be possible to write an entire encyclopedia collaboratively on an open source basis - just because the web user's did not seem to be able to write 'correct' articles....

Tuesday, November 28, 2006

UIMA - Unstructured Information Management Architecture

This morning, we were invited to a talk given by Thilo Götz from IBM about UIMA (Unstructured Information Management Framework), IBM's Framework for the Management of unstructured information that happened to take place at the department of computer linguistics.
UIMA represents (1) an architecture and (2) a software framework for the analysis of ustructured data (just for the record: structured data refers to data that has been formally structured, a.g. data within a relational database, while unstructured data e.g., refers to text in natural language, speech, images, or video data). The purpose of UIMA is to provide a modular framework that enables easy integration and reuse of data analysis modules. In general, the UIMA framework distinguishes three steps in data analysis:

(1) reading data from distinguished sources
(2) (multiple) data analysis
(3) presentation of data/results to the 'consumer'

Also it enables remote processing (and thus simple parallelization of analysis tasks). Unfortunately, at least up to now, there is no GRID support for large scale parallel execution.
Also, simple applications of UIMA, e.g. in semantic search were presented (although their approach to semantic search means: do information retrieval on unstructured data and fit the resulting data into the index of the 'semantic search engine'...)
Nevertheless, we will take a closer look at UIMA. We are planning to map the workflow of our automated semantic annotation process (see [1]) into the UIMA architecture and I will tell you about our experiences made....
UIMA is available as a free SDK, and the core Java framework is also available as open source.

References:
[1] H. Sack, J. Waitelonis: Automated Annotations of Synchronized Multimedia Presentations, in Proceedings of Mastering the Gap : From Information Extraction to Semantic Representation (MTG06 / ESWC2006), Budva, Montenegro, June 12, 2006.

Tuesday, November 21, 2006

Document Retrieval vs. Fact Retrieval - In Search for a Qualified User Interface


Today, if you are looking for information in the Web, you enter a set of keywords (query string) into a search engine and in return you will receive a list (= ordered set) of documents that are supposed to contain those keyword(s) (or their word stem). This list of documents (therefore 'document retrieval') is ordered according to the document's relevance with respect to the user's query string. 'Relevance' - at least for Google - refers to PageRank. To make it short, PageRank reflects the number of links referring to the document under consideration, each link weighted with its own relevance being adjusted by the number of total links starting at the document that contains this link (in addition with some black magic that is still under copyright restriction, see U.S. Patent 6285999).
But, is this list really what the user expects for an answer? O.k. meanwhile, we - the users - have become used to this kind of search engine interface. In fact, there exist books and courses about how to use search engines in order to get the information you want. Interesting fact is that it is the user, who has to get adapted to the search engine interface....and not vice versa.
Instead it should be the other way around. The search engine interface should get adapted to the user - and even better to each different user! But, how then should a search engine interface should look like? In fact, there are already search engines that are able to give the answer to simple questions ('What is the capital of Italy?'). But, they stil fail in answering more complex questions ('What was the reason for Galileo's house arrest?').

In real life - at least if you happen to have one - if you are in need for information, you have different possibilities to get it:

  1. If there is somebody you can ask, then ask.
  2. If there is nobody to ask, then look it up (e.g. in a book).
  3. If there is nobody to ask, and if there is no way to look it up, then think!

Let's consider the first two possibilities. Both do also have their drawbacks: Asking somebody is only helpful, if the person being asked does know the answer. (O.k., there is also the social aspect that you might get another reward just by making social contact...instead of getting the answer). If the person does not know the answer, maybe she/he knows, whom to ask or where to look it up. But we might consider this fact as being a kind of referential answer. On the other hand, even if the person does know the answer, she/he might not be able to communicate the answer. Maybe you speak different languages (not necessarely different languages in the sense of 'English' and 'Suaheli', but also consider a philosopher answering the question of an engineer...). Sometimes you have to read in between the lines to understand somebody's answer. At least, in some sense we have to 'adapt' to the way the other person is giving the answer to understand the answer.
Considering the other possibility of looking up the information, we have the same situation as if asking the www search engine. E.g., if we look up an article in an encyclopedia, we use our knowledge of how to access the encyclopedia (alphabetical order of entries, reading the article, considering links to other articles...being able to read...).
Have you realized that in both cases we have to adapt ourselves to an interface. Even when asking sombody, we have to adopt to way this person is talking to us (her/his level of expertise, background, context, language, etc.). From this point of view, adapting to the search engine interface of Google seems not to be such a bad thing at all....

If it comes to fact retrieval, the first thing to do is to understand the user's query. To understand an ordinary query (and not only a list of interconnected query keywords), natural language processing is the key (or even as they say the 'holy grail'). But even, if the query phrase can be parsed correctly, we have to consider (a) context and (b) the user's background knowledge. While the context helps to disambiguate and to find the correct meaning of the user's query, the user's background determines its level of expertise and the level of detail in which the answer is best suited for the user.

Thus, I propose that there is no such thing as 'the perfect user interface'. Anyway, different kind of interfaces might serve for different users in different situations. No matter how the interface will look like, we - the users - will adapt (because we are used to do that and we learn very quickly). Of course, if the search engine is able to identify the circumstances of the user (maybe she/he's retrieving information orally with a cell phone or the user is sitting in front of a keyboard with a huge display) the search engine may choose (according to the user's infrastructure) the suitable interface for entering the query as well as for presenting the answer...

WebMonday 2 in Jena - Aftermath


Yesterday evening the 2nd WebModay took place in Jena Intershop Tower. I thought that the number of participants that happend to come by the last time could not be surpassed (we had almost 50 people up there), but belief it or not, I counted more than 70 people this time! Lars Zapf moderated the event and we had 4 interesting speakers this evening.
For me, the most interesting talk was the presentation of Prof. Benno Stein from the Bauhaus-University Weimar about Information Retrieval and current projects. He was addressing the way how we are using the web today for retrieving information. Most current search engines are only offering 'document retrieval', i.e. after evaluating the keywords given in the user's query string the search engine presents an ordered list of documents that the user has to read in order to get the information. Instead, the more 'common' way to get information would be to ask a question and to receive an 'real' answer (= fact retrieval). I will discuss these different types of 'user interfaces' in an upcoming post. Interesting thing to mention is that Weimar is so close to Jena and both our research really seems to have some interconnections (thus, this new contact might be considered to be another WebMonday's networking success).
After that, Matthias Leonhard was giving the first part of a series of talks related to Microsoft's .NET 3.0.
Then, Ryan Orrock addressed the problem of 'localisation' and translation of applications. If translating an application into another language, simple translation of all text parts is not sufficient. There are also different units of measure to consider as well as the adaption of screen design, if texts in different languages have diferent sizes.
In the last presentation Karsten Schmidt was addressing networking with openBC/Xing, an interesting social networking tool that is supposed to make business contacts.(At least, now I know that I need some other tool to store (physicaly) my (and other people's) business cards :) ).
Even more interesting was - as always - the socializing part after the presentations. Markus Kämmerer made some photos .

Here you can find other blog articles on the 2nd WebMonday: