Showing posts with label semantic web. Show all posts
Showing posts with label semantic web. Show all posts

Tuesday, September 17, 2013

Some answers about the heady stuff of #Wikidata

I asked Emw questions as a result of his email about the migration away from the "GND main type". I am happy with the answers I received and I hope you will enjoy reading them.
Thanks,
      GerardM

At Wikidata, I contribute to discussions about properties, where I espouse using W3C recommendations and conventions from the wider Semantic Web. I'm also active in discussions about how to model molecular biology data.  Outside of Wikidata, I've been an active contributor to Wikipedia and Commons for several years.

I only had time to answer three of your questions, but I did that much pretty extensively.  The remaining questions are mostly beyond my knowledge and I don't have any well-formed opinion on them.  If you'd like, I can try to answer those questions or others next week.

My answers to your questions:

1) The GND system has been ditched. Can you explain why this is a good thing?
The GND main type property has several major problems. Deprecating that property helps us focus on better solutions for classifying knowledge on Wikidata.
Major issues with P107:
  1. With the GND main type property, "person" can mean things well beyond the common understanding of that word.  It can mean things like Coco Chanel -- i.e. 'person' as conventionally understood -- or it can mean a god, literary character, pseudonym, collective pseudonym or spirit.  The standard response to this glaring issue is "'person' is meant to generalize, don't take the term literally". That is not a sufficient solution.  If a classification system for all human knowledge considers Vishnu and Coco Chanel to be both be 'persons', that's a big problem.  Beyond giving users bizarrely unexpected query results, it means properties that should be safe to assume for any given 'person' item simply cannot be.
  2. Any item that is not a person, place, event, organization or work is classified as a "term", which contains virtually no information.  We need to be able to classify things like gravity, carbon, DNA, cancer, clarinet, Twelver Shia Islam, fashion boot, dog and potato as more than simply "terms".  One sixth of the property is kruft.
  3. Not even the GND directly uses GND main types.  The GND Ontology has a hierarchical class system and the Deutsche Nationalbibliothek -- which developed it -- uses the lowest-level, most specific GND class available for a subject.  This indicates that the GND senses the GND main types are not appropriate to use as they are with P107.
  4. The nature of P107 implies that the property is only for the highest level of classification, and that additional properties would be needed for each level in the hierarchy of classification for lower-level types. This would entail lots of unnecessary work to create and update classifications. For example, want to specifically classify Nauru? If property P107 were to persist, then you would need to add something to the effect of "main type: Place" and "subtype: Administrative unit". The problem gets drastically worse for subjects with more levels of classification, like organisms, instruments, molecules, diseases, towns, etc.
The GND system itself -- the GND Ontology -- is not the real problem.  The real problem is that P107 is a "main type" property.  In a project to structure all knowledge -- which Wikidata is -- restricting all items into a small set of types will inevitably lead to many, many classifications that are either A) too broad to be useful or B) simply incorrect.
2.  You sent an email where you asked for attention for what is to be next. Why should there be something next?
Because -- although it is complex -- the world has structure, and classes or types are a useful way to express that structure.  The lopsided debates in the Primary sorting property RFC indicate that so-called "main type" properties (sometimes also called "principal group" or "primary sorting" properties) are a bad idea.  However, that does not mean that the basic notion of grouping things into "types" or "classes" is also a bad idea.
A much better solution for classifying things is to use "type" properties recommended for the Semantic Web by the W3C -- that is, use rdf:type and rdfs:subClassOf. These properties exist in Wikidata as instance of (P31) and subclass of (P279). These properties have been part of W3C recommendations for the Semantic Web for almost a decade. They are fundamental properties used in large controlled vocabularies to structure data into knowledge.  They facilitate classification at an arbitrary granularity.  Together 'instance of' and 'subclass of' can classify all subjects and be used to determine precisely where each subject exists in the hierarchy of knowledge -- or, perhaps -- a collection of hierarchies of knowledge.

Not only do they solve those structural problems of P107 and other "main type" properties, but by being based on W3C recommendations, instance of (P31) and subclass of (P279) also make Wikidata more interoperable with the rest of the Semantic Web.
That said, deciding on properties like P31 and P279 is only the beginning of forming a better way to do classification on Wikidata.  We need a way to map the information in P107 to use P31 and P279.  That's a topic of active discussion on Wikidata.
3.  The GND is a library system then you mention upper ontologies. What is the difference, and how are they practical in the Wikidata context?
The GND (Gemeinsame Normdatei) authority file is used as a library classification system, but it's based on the GND Ontology.  The ontology has a hierarchy of high-level entities and sub-classes.  The P107 property is based on those so called "high-level entities", which were called "main types" in Wikidata as shorthand.  The main GND types are person, place, event, organization, work, term or "undifferentiated person".  These main types are fine as a way to classify items of general interest in a large library, but they're much too small to form a sound basis for a classification system for all human knowledge. 

That's what upper ontologies are for.  An upper ontology is a way to have standard vocabulary about high-level entities in our world.  The idea is to formalize these very general concepts in a way that captures the richness of human language while also being precise enough to be machine-understandable. 

For example, the Suggested Upper Merged Ontology (SUMO) sets a class "entity" as the most general type of thing -- everything is an "entity".  From there, SUMO classifies things in the world as either "physical" or "abstract".  "Physical" things can be "objects" or "processes".  "Abstract" things include so-called "set-classes", "propositions", "quantities" and "attributes".  (More information on SUMO is available in Towards a standard upper ontology.)
There are several other upper ontologies available, like BFO and UMBEL.  I am not an expert in ontologies, and I have not learned enough about each of them to make an informed statement on their advantages and disadvantages.  However, because they seem to offer unifying terminology for different domains of knowledge, upper ontologies strike me as something worth consideration by the Wikidata community.

Saturday, March 31, 2012

#Wikidata, the interview

The press release is out, the mailing lists are full of it. But what is this brand new Wikidata project about. Who better to ask but Lydia Pintscher and Daniel Kintzler. Lydia does "community communications" for the Wikidata project and Daniel has been involved in many data related projects including OmegaWiki, the first Wikidata iteration.

Wikidata is still brand new; it does not have its own logo however it is ambitious and I expect that it will improve the quality and the consistency of the data used in MediaWiki projects everywhere. Enjoy the answers to the ten questions answered by Lydia and Daniel.
Thanks,
      GerardM

What is it that Wikidata hopes to achieve?
The Wikidata project aims to bring structured data to Wikipedia with a central knowledge base. This knowledge base will be accessible for all Wikipedias as well as 3rd parties who would like to make use of the data in it.

When you express this in REALLY simple language, what is the "take home" message of Wikidata?
We are creating a central place where data can be stored. This could for example be something like the name of a famous person together with the birthdate of that person (as well as a source for that statement). Each Wikipedia (and others) will then be able to access this information and integrate it in infoboxes for example. If needed this data can then be updated in one place instead of several. There is more to it but this is the really simple and short version. The Wikidata FAQ has more details.

Structured data is very much like illustrations. The same data can be used over and over again. Will there be a single place to update for everywhere where it is used?
Yes. There will be a central place but we are also working on integrating this in the editing process in the individual Wikipedias.

This project is organised and funded by the German chapter. Will this ensure that the data can be used in many languages?
The German chapter is indeed organising it. Funding is coming from Google Inc., AI2 and the Moore Foundation. One of the main points of Wikidata is that it will no longer be necessary to have redundant facts in different Wikipedias. For example it is not really necessary to have the length of Route 66 in each language’s article about it. It should be enough to store it once, including sources for the statement, and then use it in all of them. In the end the editor will be free to decide if he or she wants to use that particular fact from Wikidata or not. This should be especially helpful for smaller
Wikipedias who can then make better use of the work of larger Wikipedias.
In order to make the data useful in pages written in different languages, we of course have to provide a way to supply information in different languages. This is described in more detail below.

In new Wikipedias a lot of time is spend in localising info boxes. Will Wikidata make this easier?
The template for the infobox still has to be created by the respective Wikipedia community -- but filling the infoboxes would be much easier! Wikidata could be used to get the parameters for the infobox templates automatically, so that they do not need to be provided by the Wikipedians.
    Lydia

Will it be possible to associate the data with labels in many languages? 
Each entity record can have a label, a description (or definition) and some aliases in every language. Not only that: each language version of the label can have additional information like pronunciation attached. For example, the record representing the city of Vienna may have the label “Vienna” in English and “Wien” in German, with the respective pronunciations attached (/viːˈɛnə/ resp [viːn]).

Many people who are part of the project have a Semantic MediaWiki background. How does this affect the Wikidata project? 
Wikidata will profit from the team’s experience with Semantic
MediaWiki in two ways: they know what worked for SMW, and they know what caused problems. We plan to be compatible to the classic SMW in some areas: for instance, we plan to re-use SMW plugins for showing query results. On the other hand, we will use a data model and storage mechanisms that are more suitable to the needs of a data set of the style and size of Wikipedia.

To what extend will Wikidata data be ready to be expressed in a "semantic" way and if so what are the benefits?
Wikidata will only express very limited semantics, following the linked data paradigm rather than trying to be a true semantic web application. However, Wikidata will support output in RDF or the resource description framework, albeit relying on vocabularies of limited expressiveness, such as SKOS.

The DBpedia project extracts data from the Wikipedias. To a large extend this is the same data Wikidata could host. Is there a vision on how Wikidata and DBpedia will coexist?
If and when all structured data currently in Wikipedia is maintained within Wikidata, the extraction part of DBpedia will no longer be necessary. However, a large part of DBpedia’s value lies in the
mapping and linking of this information to standard vocabularies and data sets, as well as maintaining a Wikipedia-specific topic ontology. These things are and will remain very valuable to the linked data
community.

You worked on the first Wikidata iteration; OmegaWiki. What is the biggest difference ?
The idea of Wikidata is quite similar to OmegaWiki - it’s no coincidence that the software project that OmegaWiki was originally based on was also called “Wikidata”. But we have moved on since the
original experiments in 2005: The data model has become a bit more flexible, to accommodate the
complexity of the data we find in infoboxes: for a single property, it will be possible to supply several values from different sources, as well as qualifiers like the level of accuracy. For instance, the
length of the river Rhine could be given as 1232 km (with an accuracy of 1km) citing the Dutch Rijkswaterstaat as of 2011, and with 1320 km according to Knaurs Lexikon of 1932. The latter value could be marked as deprecated and annotated with the explanation that this number was likely a typographical error, misrepresenting earlier measurements of 1230 km.
This level of depth of information is not easily possible with the old Wikidata approach or the classic Semantic MediaWiki. It is however required in order to reach the level of quality and transparency
Wikipedia aims for. This is one of the reasons the Wikidata project decided to implement the data model and representation from scratch.
     Daniel

Thursday, April 15, 2010

#DBpedia 3.5 has been released

#Wikipedia aims for its content to be used. Reading it, mashing it, analysing it, republishing it whatever.. DBpedia is an amazing example of what can be done with all the data gathered. To quote from the 3.5 release announcement:

The new DBpedia knowledge base describes more than 3.4 million things, out of which 1.47 million are classified in a consistent ontology, including 312,000 persons, 413,000 places, 94,000 music albums, 49,000 films, 15,000 video games, 140,000 organizations, 146,000 species and 4,600 diseases. The DBpedia data set features labels and abstracts for these 3.2 million things in up to 92 different languages; 1,460,000 links to images and 5,543,000 links to external web pages; 4,887,000 external links into other RDF datasets, 565,000 Wikipedia categories, and 75,000 YAGO categories. The DBpedia knowledge base altogether consists of over 1 billion pieces of information (RDF triples) out of which 257 million were extracted from the English edition of Wikipedia and 766 million were extracted from other language editions.

All this data is there to be used. That is the whole point to DBpedia. It makes all this data available in a format that makes it easy to mash, to analyse, to verify. The key thing is that through DBpedia Wikipedia is linked to many data sources.
  • we can project geo data on maps
  • we can compare / verify data with other sources
  • we can improve the consistency of our data
  • we can compare the data between the different language versions
  • we can link our illustrations to the GLAM that hold the original
  • we can link our sources to public resources where you can read the original text
  • we can link our sources to the libraries where you can find a source
DBpedia is a tool that exists today, a tool that wants to be used. It is the kind of tool that helps us out of our isolation and provides us with a niche in the wider data world
Thanks,
      GerardM

Wednesday, November 01, 2006

The semantic web to the rescue ?

The reporting on the Internet Governance Forum in Athens is mighty interesting. Yet another nice article on the BBC website with some thought provoking ideas.

I find it really interesting that spoken languages are considered. However, practically at this stage the Internet is very much oriented towards written languages and, given the amount of stumbling blocks that exist to integrate languages other than the ones using a Latin script, I am afraid this is just a red herring.

I was also amused to see that the semantic web was brought to the fore as one solution to the problem of linguistic diversity. Yes, it is intended to be understood by computers. Computers are used by people and the semantic web is decidedly English. This raises the question how this computer that apparently understands English communicates to its user who does not.

There are however some great things to be said about the semantic web; first of all the terms used should be unambiguous. This in turn means that it should be possible to translate it to other languages than English. This is a challenge that we face in WiktionaryZ. We are able to have semantic relations and our semantic relations do translate to the language of the User Interface. So when the terms have been translated, in WiktionaryZ relations can be understood not only but also by computers.

This is a good moment for a disclaimer; WiktionaryZ is pre-alpha software. Many of the issues that have been tackled in the development of the semantic web we have not considered let alone touched. We hope / expect that we will be allowed to stand on "the shoulders of giants".

Thanks,
GerardM