Friday, October 18, 2013

#SignWriting, #sign languages - an #Interview with Valerie II

Half of the people who sign are not deaf. Therefore the number of people who benefit from a sign language that can be written is more than just the number of people who are deaf and sign.

Thanks to SignWriting, sign languages can move on from only having an "oral tradition". I have been privileged to witness as the SignWriting community moves slowly but surely ahead in gaining recognition by making their languages and their cultures equal to any other language in the age of Internet.

Knowing Valerie is an inspiration. She is a real mover and shaker for so many people. I rate her as highly as Jimmy Wales. Anyway, these are the questions I put to her. Enjoy!
Thanks,
       GerardM

How do you explain what it means when a language cannot be written?
If I understand it correctly, most of the world's languages do not have a written form. All languages CAN be written. But most languages are not written.

Sign languages are now written languages!  ;-))

But it takes effort by people who know their languages to want to develop a way to write it and lots of languages, for example, in Africa and Asia, may not have the political position, nor the funds, to invest in the development.
How many languages are written in the SignWriting Script?
That is also hard to say, but we estimate that small groups of people are writing their sign languages in around 40 countries…based on real written literature and also by word of mouth…we have definite proof for many, and some proof for some…. Some sign languages have 100s if not 1000s of written documents - ASL and DGS are two examples
Can you recognise what sign language it is from a written text?
Yes, between written ASL and written DGS (German Sign Language) we can definitely see a difference immediately - Other sign languages, not so much yet, because we do not have that much experience comparing written documents - as each sign language has more and more literature, it will be easier to recognize differences in the sign language literature - a lot has to do with the choice of style of writing… German Sign Language is written with more mouth movements than our ASL written documents… so one can see which language it is quickly
What are the most active sign languages (that are writing with SignWriting)?
American Sign Language, Argentine Sign Language, Brazilian Sign Language, Catalan Sign Language, Czech Sign Language, French-Canadian Sign Language, French-Belgian Sign Language, French-Swiss Sign Language, French Sign Language, Flemish Sign Language, German Sign Language, German Swiss Sign Language, Italian Sign Language, Jordanian Sign Language, Korean Sign Language, Maltese Sign Language, Nicaraguan Sign Language, Norwegian Sign Language, Polish Sign Language, Portuguese Sign Language, Saudi Arabian Sign Language, Slovenian Sign Language, Spanish Sign Language, Tunisian Sign Language and more…

The use of SignWriting is growing rapidly. How do you know about how it develops?

Through the internet, in different ways. And through individuals sending me documents. And through mentions on the SignWriting List. Some write documents publicly in SignPuddle. Some get their school degrees - Ph.Ds and Master Degree theses are posted written on SignWriting or using SignWriting… Papers are presented about SignWriting and they are listed in publications, and occasionally people write to me privately or join the SignWriting List or Facebook or Twitter… but I actually do not know how many people use SignWriting and I never will, because it is free on the Internet and the way it spreads is like it has a life of its own
Can SignWriting be used on mobile phones or is there an app for that…
Yes, we have two apps for the iPhone…one from Germany and one from California:
SignWriting App from California, by Jake Chasan (age 16 ;-)
Signbook App from Germany, by Lasse Schneider, University of Hamburg
Are there many schools where they teach SignWriting
SignWriting is spread freely on the internet, so lots of people learn SignWriting on their computers, but there are schools with official courses too, such as

  • Osnabrück School for the Deaf in Germany
  • A school in French-Canada (Quebec)
  • Schools in French-Belgium
  • A school in Flemish-Belgium (Brussells)
  • A school in Poland
  • A University in the Czech Republic
  • A Catholic School for the Deaf in Slovenia
  • Bible Translators teach SignWriting in Madrid
  • A university in Barcelona (Catalan Sign Language)
  • ASL classes in a hearing high school in Tucson, Arizona
  • ASL classes at San Diego Mesa Community College in San Diego, California
  • ASL classes at UCSD, University of California San Diego
  • Courses on SignWriting are taught throughout Brazil, by Libras Escritas, also online
  • A School for Deaf Children in Brazil: Teacher Sonia Messerschimidt
  • Santa Maria-Rio Grande do sul - Brasil , Escola Estadual de Educação Especial
  • Letras LIBRAS , UFSC Florianópolis, (SignWriting is an integral part of the curriculum)
  • the list goes on and on ... 
How hard would it be to have the pupils at these schools write two articles a month ... How many Wikipedias could be started that way…
Just as soon as Steve Slevinski and Yair Rand are ready for us to move Nancy Romero's 37 articles, and Adam's 2 articles and Charles Butler's 1 article, from the Wikimedia Labs to the Incubator, and we have tested the ASL User Interface, and we have tested the new features like linking and selecting text, and when Steve has completed the new SignWriting Editor program that will make it possible to write articles directly on the Incubator site, then of course we can ask teachers and students to test our new software and start writing articles… our software isn't ready yet though. 
That is why Nancy Romero, our most prolific English-to-ASL translator and SignWriter, is writing enough articles to lay a foundation so we can get started - we could ask students to write the articles over in SignPuddle Online, and then we can move those articles over to the Incubator for them - but until the Editor software is completed it wouldn't be as much fun for the students as it will be later - that day is coming and it is an excellent idea for the future.
Why is Wikipedia strategically important for getting more people to know about SignWriting.
I consider it VERY important because it provides us literature for readers to read written in ASL that are not children's stories or religious literature. Ironically we have plenty of children's stories written in ASL, including Cat in the Hat, Goldilocks, Snow White and others…and Nancy Romero has written close to the entire New Testament in ASL based on the New Living Translation, and another Deaf church has also written much of the Bible - but we are really grateful to have articles to read in Wikipedia that are general non-fiction - educational, historical, scientific.
Wikipedias are also important because they will encourage others to write articles in ASL which will indirectly teach people how to write ASL. Another wonderful indirect result will be an added respect from the general public. Surprised visitors will realize that ASL and other sign languages can now be written -
Val ;-)

Reasonator gets "instance of" "#human"

Reasonator, the tool by Magnus now appreciates that a person is a "human". An increasing number of items in Wikidata identify a person as a human. As a consequence, all the tools that rely on appreciating that an item is about a human need to adapt. When you are aware of tools that still need to adapt, pinging the author of the tool is a reasonable approach :)


I selected Herman Melville for the illustration; he is mentioned today as the author of Moby Dick on the main page of the English language Wikipedia.
Thanks,
         GerardM

#Statistics for #Wikidata

Finally some up to date and useful statistics for Wikidata. Magnus Manske brings us these wonderful statistics that are up to date and actually useful. To appreciate the one above, you have to know that there are many items who are linked to outside sources like VIAF. Everything above 5 statements has probably contributions made by volunteers.

When an item has no label in a language, you cannot search for it. If anything this is what has me worried. To be useful in a language you have to be able to find an item. Given that there are 280+ languages represented in Wikidata, this is probably the biggest challenge facing Wikidata.
Thanks,
      GerardM

Thursday, October 17, 2013

#VIAF, sources and the battle of the sexes continued ..

In the English language #Wikipedia there are more articles about men. In #Wikidata there are more persons known than in VIAF and both are aware of more men as well.

When you research Wikidata you may learn the gender ratio of men and women in any Wikipedia. As more statements are added, the data will become more precise and more articles will become known to be about "humans".

VIAF is maintained by the OCLC and it brings together the information of many libraries from around the world. It links people; authors and subjects to books. Wikipedia does to a certain extend the same thing by linking sources to facts known in articles. Many of these sources are books or periodical publications and identified by an ISBN or ISSN number.


The notes, references and further reading used in Wikipedia articles do refer to books and periodical publications and often provide the ISBN and ISSN numbers as well. They may even refer you to your library.

Wikidata allows for adding sources to its statements. Typically there are no real sources for any statement on any item. Sadly this is used as an excuse for not using Wikidata.

When all the Wikipedias are mined for its sources and when they are added to Wikidata as sources for a subject, Wikidata will become a useful tool for the library world. For people who love to add sources, it becomes much easier to add sources to statements but also to verify the veracity of statements.

As to the battle of the sexes, when you care about more articles about women, you have to write them. There are so many notable humans who do not have an article yet.
Thanks,
       GerardM

#SignWriting, #sign languages - an #Interview with Valerie I

The biggest challenge for a #language to gain permanence is to be written. Many languages became written languages by adopting an existing script. Sign languages are fundamentally different and therefore there was no script to adopt.
Valerie Sutton, a ballerina, developed a method to register dance movements. Linguists who researched sign languages asked Valerie if this could be applied to sign languages. Many iterations later, SignWriting became an ISO recognised script, it is known to be used by at least 40 sign languages that all may gain their Wikipedia.

When I asked Valerie to answer ten questions, her response was to ask the SignWriting community for their opinion. This, the first part, explains the need for sign languages. As it is the most often asked question about sign languages it deserves a full response.
Enjoy,
      GerardM

Some people say "Why do they not use English?" .. How different is a sign language from a spoken language?
A lot of Deaf people DO use English … ;-)

Deaf people who use a sign language as their primary daily language, also use English as their second language. But, they cannot hear their second language, English, and American Sign Language (ASL) and other sign languages are rich languages that give a deep communication that is more profound than speaking a second language you struggle to hear, or cannot hear at all… Lip reading does not give all the sounds made on the lips, and many conversations have to be guessed at most of the time… They say that at best, lip reading gives 30% understanding and everything else is guessed at….

So if a signing Deaf person, whose primary language is American Sign Language, lives is in the United States or English-speaking Canada, they have to get around, and they learn English to get by…

Just as I learned Danish when I lived in Denmark. Learning a second language is a requirement and you do your best… But there was one difference for me - I can HEAR Danish, which was my second language years ago when I lived in Denmark… I would have found it much harder to learn Danish if I were Deaf and could not hear it…And truth be told, if I really wanted a deep and profound communication, I would always migrate back to my native language English. My second language did not give me the true communication of my native tongue.

So the question is really not "why one language is better or easier that the other?" but instead a realization that deafness creates a barrier to learning spoken languages, and that first and second languages are different experiences… Deaf people are not choosing one language over the other, but instead managing the best they can between the hearing and Deaf communities.

Both signed and spoken languages are good languages and should be equally respected. And it is my feeling that everyone should learn another language if they can. Hearing people who speak English as their native language oftentimes enjoy learning American Sign Language, but no one asks them why they just don't use English?! (at least I hope not ;-)

In school we are asked to learn a foreign language…so why not learn American Sign Language or other signed languages?

And in return, Deaf people spend most of their lives learning spoken languages to the best of their ability and I give them my utmost respect for the hard work I know they must go through everyday...

Some Deaf people are born into Deaf families. Deaf children in Deaf signing families have a native language, sign language, from the beginning, so their language development is early and considered the same as a hearing child's language development. Spoken languages are a "second language" to everyone in the Deaf family. So they do not feel "different" than their parents or siblings.

Deaf children born into hearing families sometimes have a harder time, because oftentimes the family doesn't even know the child is deaf until later, and so language development may not start early, and also they are different than their own parents and family members.

So native signing Deaf people have their own native language, a sign language, and yes, the grammar and structure of American Sign Language, for example, is quite different than the grammar and structure of English. Verbs are conjugated differently, adverbs and adjectives are in different positions in the sentence, and there are elements of American Sign Language that are much more sophisticated than in English, or at least expressed very differently, and so oftentimes there is not a real way to translate between the two languages that is a true "match"… what takes a paragraph in English can be expressed with a short phrase in ASL, and vice versa… Some say that the grammar of ASL is closer to Russian or Spanish than it is to English...

That is why SignWriting is important. When both languages can be written, both languages can be compared, and understood better…

Val ;-)

Monday, October 14, 2013

Battle of the sexes, the #Wikidata story III

#Freebase is a rich resource. It boasts to have 39,296,842 topics and much of it is linked to #Wikipedia. Freebase has been around for a considerable amount of time and, there are great applications that make use of the data.

One of these applications informs you if a name is typically male or female. I would not expect the name Roxy to be used by men but it is.

The information for many of the people involved is linked to Wikipedia. Consequently it is relatively easy about the sex of someone like Roxy Walters and update Wikidata. At this time no statements have been made for Roxy Walters.

The only potential problem is that Freebase is using the CC-by-sa license and that makes it incompatible with Wikidata. Given that you cannot license facts, Google may decide either to change the license for Freebase or ignore when Freebase is used to update Wikidata.
Thanks,
       GerardM

The distribution of page views for #Wikipedia

More people are reading Wikipedia than ever before. The year on year growth for all Wikipedias is 35% and the prognosis for October is that this will at least continue.

At the bottom of the page view statistics you will find the distribution of the page views grouped by the size of the Wikipedia.

What is really interesting is that this last year the Wikipedias that are not in the top ten by size are growing faster than the 10 biggest Wikipedias. This change is around 2%.

So far the most attention has been going to the English Wikipedia. It has grown 28% YoY. When you compare this to the growth of all Wikipedias combined (35% YoY), it is obvious that the numbers suggest to put more effort in the other projects for best effect.

The most obvious and cheapest effort the WMF could take is by making sure the developers take their responsibility for the internationalisation of the software they write. The backlog of support requests is here.
Thanks,
     GerardM

Battle of the sexes, the #Wikidata story II

#Wikidata is the most obvious place to study the distribution of articles among the sexes. There is however a problem; many "items" representing people do not indicate that they are "human" or their sex.

One approach to this problem to this problem has been indicated by Markus Krötzsch. In a blog post he describes how he determines the sex of a person based on a first name. He produced a long list of first names and the probable sex of people with that name.

With this approach it is possible to determine the overall distribution of the sexes within Wikidata. It should also be possible to look at the distribution of the sexes within the different Wikipedias because of the interwiki links. When first names are known in the Latin script, determining the first names in different scripts and thereby expand the reach of this process should be feasible as well.

As this algorithm works pretty well for research, it could be the basis for a tool. When someone is shown the first paragraph of an article with a yes/no button it should be really easy to quickly add a lot of information to Wikidata.
Thanks,
      GerardM

Sunday, October 13, 2013

What #source for #Wikidata

When statements are sourced on #Wikidata, it should be obvious; there is no single language that is to be preferred as a source. Philosophically the language best suited to provide a source is a source that is closest to the events or the persons involved.

When the subjects are historic like the "Sanjak of Bosnia", many of the facts involved are not clear; the family name of the first sanjak-bey may be either Hranić or Pavlović. It matters in that the understanding of history may depend on it.

One reason why there are so few sources in Wikidata is because it is so cumbersome. The sources and their author have to have a Wikidata item in their own right.. It would help if just an ISBN or ISNN number would suffice..

Reproducing a family tree for Mr Isa-Beg Isaković is possible but it does not make sense to do so without information I trust. With a subject related to Bosnia and the Ottoman empire, I prefer to leave this to people who are knowledgeable.

What I can do is build the succession of all the sanjak-beys that ruled the Sanjak of Bosnia. My source would be the English Wikipedia article and it does refer to several red links. Including these will help link potential articles into this framework and in a way it can make Wikidata a source for Wikipedians.
Thanks,
      Gerard

Missing #statistics in #Commons and #Wikidata

Both Commons and Wikidata provide a service to other Wikimedia projects. However, both the media files and the data can be used outside these projects as well.

Commons is where most of the media files are stored. Currently there are 18,924,867 files available and the project is advertised as "a database of freely usable media files to which anyone can contribute". When you look at the statistics for Commons, apparently the only activity that is measured is about the contribution to Commons. There is nothing about its use. 

Wikidata aims to create a free knowledge base about the world that can be read and edited by humans and machines alike. At this time there are 13,538,727 items that anyone can edit. People increasingly do edit these items but it is not known how often these items are used. The statistics for Wikidata, like the Commons data are only about the production and maintenance of the project. 

The point of any and all of the Wikimedia projects and the information they contain is that they are to be used. The success is dependent on their usage. The existing statistics while interesting do not measure the success of these projects where it counts. The usage. The question how do THEY use it with/without you is a most relevant question and, we do not begin to understand the use and the re-use of our data when we do not measure it, when we do not compile statistics about it.
Thanks,
      GerardM

Vo Nguyen Giap, a general and politician from #Vietnam II

#Wikipedia is a collaborative project and, this was demonstrated really well on the article for General Vo Nguyen Giap. A template was added with information about him. The information was made available from #Wikidata. An Occitan Wikipedian added text to the article and, it is no longer a stub.

Included in the text is a red link to the Communist Party of Vietnam and this red link showed as an identifier in the infobox of Mr Giap. The red link in the text is effectively the translation in Occitan for Communist Party of Vietnam. When this is copy pasted as the Occitan label in Wikidata, it becomes available as a red link in the info box as well. 

What is really cool is that this red link now also shows on the article of Ho Chi Minh.
Thanks,
      GerardM

Saturday, October 12, 2013

John W. Johnston, a #Virginia Senator

Mr Johnston was also a USA senator. But #Wikidata was not yet aware of the fact that he was also a member of the Virginia Senate. It did say so on the info-box.

To get that fact in, I had to create a new item for "Virginia senator". The number of people who have been in this senate or any of the other state senates must be big and most of them are of a questionable notability. But then again both Wikipedia and Wikidata can take them in as much as politicians from other countries.

When you Google for Mr Johnston, you can find the fine picture that is to the right. It is in Commons and the original can be found in the Library of Congress.

I looked at Mr Johnston because today the article about him featured on the main page of the English Wikipedia. I wanted to know how complete the information about a featured article is on Wikidata.
Thanks,
      GerardM

Thursday, October 10, 2013

Abdallah II, six times Sultan of #Morocco

When the title of Sultan is supposed to be hereditary, calling the title yours six times in a row is special. Abdallah of Morocco had the dubious pleasure to fight his many brothers to finally gain the upper hand.

On the French Wikipedia, they have visualised this nicely. This  information is available on Wikidata using qualifiers. As the use of qualifiers to indicate who preceded and succeeded as Sultan is not universally liked, the question is very much what could serve as an alternative?
Thanks,
      GerardM


Wednesday, October 09, 2013

Even more heady stuff about #Wikidata and ontologies

In an interview you ask questions and the answers are not yours. When you are as lucky as I have been, you get food for thought but not necessarily answers that close an argument. Given that it is my blog and I am commenting on Wikidata anyway, I decided to add some of my own thoughts as well.

Ontologies in their structure are representations of the world. Wikidata is used by over 280 languages and these represent many more cultures. When one upper ontology finds adoption in Wikidata it will not fit well with many of them. A good example is the conversion following the "main type (GND)" "person" controversy; it has been decided that this will become "instance of" "human" in stead. This means that every person will be identified as a specimen of the species Homo sapiens. This may be technically correct but I am sure that many people will feel alienated as a result.

For me it is very important to keep in mind what Wikidata is there for. It is first and foremost a repository of information that is to be used. The considerations of other data repositories are very much secondary to that. One key objective is to use the information as used in info-boxes. The box used for Ronald Reagan serves well to illustrate a few points. It is indicated that he was both a governor of California and a president of the United States. For both these offices, a start and stop date and a predecessor and successor are indicated. In a classic semantic annotation, each of these dates and other office holders would be on their own. As a consequence it would not be possible to tell what date applies to what and who preceded or succeeded to what office. This information can only be given in a relation and qualifiers are the tool that is available.

In all this "rigidity" is important. Ronald was an actor before he was a governor or a president. This rigidity is however dependent on the community of Wikidata; it is and remains a wiki. The logic associated with an actor or a politician should apply both to Arnold and Ronald in equal measure.

When it is indicated that somebody is a "human", it is obvious that he is born, may die and has a sex. These are obvious inferences and they can be suggested as easy and obvious qualifiers and give them a preferred order of presence. As they are easy and obvious, it should be possible to map information consistently to other resources like DBpedia.

Given that explicit links exist between Wikidata and an increasing number of external resources, it seems obvious that people will find a way to compare data. The big advantage for Wikidata is that it has an explicit purpose that will ensure that its data is going to be used. It is going to be used by the biggest possible public because the aim is to share the sum of all knowledge.
Thanks,
       GerardM

Tuesday, October 08, 2013

Hassanal Bolkiah, the Sultan of #Brunei

The #Commonweath has a new repository of images. It includes many interesting images like the picture to the right of the 29th and current Sultan and Yang Di-Pertuan of Brunei.

There are many pictures available and partcularly when they are used for educational use, you should have a look. Sadly, for use in Wikipedia these pictures are off limit as they are not freely licensed.

The way it works is, 
  • you create yourself a profile
  • you select the pictures you are interested in
  • you indicate what you will use the pictures for
  • you wait for permission to use them
  • you download the picture
  • you add an attribution
Compare this with the pictures of Mr Hassanal Bolkiah on Commons; 
  • this is the category of pictures of the Sultan
  • you select one of the pictures like this one..
  • you download the picture
  • you add an attribution
Thanks,
     GerardM

Monday, October 07, 2013

Vo Nguyen Giap, a general and politician from #Vietnam

Vo Nguyen Giap died on October 4th. He was probably one of the most successful generals of the twentieth century. He was instrumental in defeating the armies of both France and the United States.

When Mr Giap died, it became international news and as a consequence new articles were written in several Wikipedias. New statements were added for Mr Giap in Wikidata.

The proof of the pudding is also in finding what the Occitan Wikipedia makes of all this information. So I created a new article their as well and added the link in Wikidata to make it work.

I am pleased with the result. To make it look even better, I could add even more information about Mr Giap for instance him having been the minister of defence for Vietnam.
Thanks,
      GerardM


More heady stuff about #Wikidata and ontologies

When I asked Emw many questions, I received only three answers. Having an opinion on Wikidata and expressing it is hard. It feels very much like exploring new frontiers. The questions did not go away so I asked them again and, I am mighty pleased that Antoine Isaac was willing to provide me with some answers. Antoine does not do Wikidata, he is a/the scientific coordinator at Europeana.eu. He wrote an email to the Wikidata mailing list that got me interested in asking him my questions. I hope we will collaborate for our GLAM data well with Antoine and Europeana.

I exchanged several emails with Antoine but the answers do stand on their own. I will react in a follow up blog post. I hope you appreciate what Antoine has to say as much as I do.
Thanks,
     GerardM

The Wikipedia article calls upper level ontologies "political" why should Wikidata be interested in any of that?
I wouldn't call them 'political', that's a bit far-stretched. Indeed ULOs embody some abstract considerations on how to represent the world. And as said before, this can backfire as soon as you consider an open web of data, where different representations may well co-exist. Considering a single upper-level ontology as the guiding principle for everything is dangerous. But it is true that they provide valuable bodies of knowledge to re-use, and it may make sense to re-use them for specific domains (e.g. biology, geography) that could be compatible as a whole with the approach of one ULO.
Does it not make sense to group statements together as qualifiers as part of a statement (like it is done for office held [1]) ?
I like qualifiers and dislike them at the same time!On the one end, it's good to have some meta-metadata about the provenance of a statement (who made/endorsed it), or its scope (e.g. the time it applies). In your example, this is the "start date" and "end date". This practices actually fits what is happening in the RDF / Semantic Web area, where a lot of work is being done about Provenance, and many people use quad stores with 'named graphs' instead of just triple stores. 
On the other end, I am anxious about qualifiers being used with other semantics that "here's some info about a statement". In your example, "preceded by" and "succeeded by" are more difficult to interpret in this sense:  (compare with property P580 that mentions explicitly "statement"). I mean, it is possible to interpret your qualifiers as data on statement. But I really feel that people (and you?) will understand it as, say "Te Rata is the predecessor of Korokī Mahuta" and not "The statement 'Te Rata held the office of Maori Monarch' is the predecessor of the statement 'Korokī Mahuta held the office of Maori Monarch'". Which should be the right thing to do (I mean, the one compatible with the "start date" semantics).  
Note that what you told in the other email is the kind of use of qualifiers that would worry me: it is the person who has a birth/death date or a sex, not the statement 'is a person'.Of course one could see the birth and death date to influence of the 'date of validity of the statement "is a person"'. But still we'd be talking about two different things, from a knowledge representation perspective. 
Note also (just for the fun of refering to upper-level ontologies and their dangers) that for some ULOs 'is a person' doesn't have a begin and end date. Being a person is 'rigid' i.e. it must stick to the subject forever. You can't have been a person once and then cease to be a person. Even if you die you're still a person...And don't tell me that rigidity foresees the some ontologies may have Person an an anti-rigid property. This may probably not be the choice made in your favourite ULO. Unless it's one that addresses both reality and beliefs as two possible sides of a same property. But then, good luck re-using it!!!
DBpedia does not have qualifiers, will this impact their ability to use data from Wikidata?
I can't really speak for DBpedia. But I'd say that if qualifiers are used in a way that is both consistent and compatible with the understanding of 'named graphs' in RDF, then they might be interested.
As we map Wikidata items to the content in other repositories, what do we need to compare the data from these repositories
Which repositories are you talking about? Which content? Are 'repositories' knowledge bases, like OCLC or Europeana in the book/artwork domain? Is 'content' 'data'? If yes, then what is required to establish correspondences is hard work looking at the fields and seeing how they correspond. This may be of course alleviated if all (wikidata included) look at what's already happening and try to minimize the risk of coming with data models that are too indiosyncratic. (this is why something like RDF is quite useful!)
When differences in content between repositories are found is there a standard method to harmonise the content
Again assuming a reading of 'repositories' and 'content' as above. I don't think there is a standard method. What you can prey for (and of course as designers of data repositories, we are somehow in the position of making it happen!) is that all repositories keep track of as many unambiguous identifiers they can keep track of (e.g. ISBNs for books) which would help automatic reconciliation. Otherwise make sure the data on the content (where 'content'='the object in the real world') is as complete as possible. For works of art that would mean that comparisons can be made on titles, creators, dates and place of creation, etc.

Sunday, October 06, 2013

500,000 pageviews .. the love for #Statistics

What I love about statistics is that you can interpret them in so many ways. Thanks to Lydia I passed the 500,000 pageviews mark. This is a nice round number but it does not mean that much. What I would like to know is the number of people who read my blogposts. My blog is syndicated on for instance en.planet.wikimedia.org and consequently there are many, many more people who have at least noticed that I blogged when I did.

The most popular article was #DBpedia and #Wikidata, the person I admire most is Valerie.
Thanks,
        GerardM

Saturday, October 05, 2013

10 #Wikidata questions for Lydia

Lydia is the new project leader for Wikidata. She has already done a great job communicating for the Wikidata project and, with much of Wikidata well established, communication will increasingly be the critical success factor. The Wikidata team works together well so I think she will do really well. Enough reasons to ask her some questions. I hope you will enjoy them as much as her answers.
Thanks,
     GerardM

Wikidata exists, it is being used. Do we know how much it is used?
Unfortunately not. Of course we get to see some great uses but not the whole picture. I love seeing all the different ways the Wikipedias are already now using Wikidata that I know about. For example the Occitan Wikipedia building whole infoboxes based on data from Wikidata. Or the English Wikipedia that compares its local IMDB identifiers against those in Wikidata and puts pages into a maintenance category if they do not match so a human can look into it and fix it. But this is only inside the big Wikimedia projects. At the same time people are building 3rd-party tools that use data from Wikidata. The most recent and very impressive one is the Wikidata tempo-spatial display. We are seeing more and more of these pop up and I am looking forward to what other useful, cool and even crazy things people will come up with.
You want people to trust Wikidata data. How do you envision this to work?
Wikidata has just started and we're seeing people add a lot of data to it for all the world to use. This is fantastic. At the same time we need to make sure that the data in Wikidata is reasonably trustworthy. I say reasonably trustworthy because in the end this is just like Wikipedia. It's not 100% perfect but we're all making a huge effort to keep it as accurate and correct as possible and people have come to rely on this. Wikidata is in a bit of a better situation here though than Wikipedia for a few reasons. First of all since the data is structured and machine-readable it is much easier for a computer to find inconsistencies and alert an editor about them. One example would be that a political office is said to be held by X but X is said to be an animal. Now there were probably a few cases where a political office was held by an animal but in general things like this should be flagged for an editor to reexamine it. And there are many such cases that could be checked. Here you find a few such checks that are already in place. The other advantage Wikidata has is that it will be much easier to verify a given data point against an external source once that is given in Wikidata. And the third advantage Wikidata has is that it will be watched by potentially a lot more people. Changes in Wikidata show up on the watchlists of all editors who are watching the corresponding page on their Wikipedia as well as the recent changes of that Wikipedia. So all in all we are in a pretty good position. However we are not there yet. More tools will need to be developed, existing ones improved and most importantly people will need to spend time adding sources to statements in Wikidata.
When you talk about the user experience of Wikidata, what are the limits of this user experience as far as you are involved?
I want Wikidata to be a joy to use and I want it to be easy to use. At the same time experienced editors need to be able to navigate the site and complete their tasks quickly and efficiently. We will have to find the right balance there over and over again. The same thing goes for all the missing features we still have to develop or roll out - queries and the numbers datatype for example. I will put more emphasis on user experience but moving the project forward on a feature level is also very important.
One obvious target for Wikidata is to include all the information contained in info boxes. How far off are we before Wikidata can service most subjects that have info boxes.
I think the big missing piece of the puzzle is the numbers datatype. We're not too far from rolling out a first version of it. Once we are able to also deal with units we are well on our way to that target.
You want people to better understand Wikidata. Is that not even more complicated than understanding templates and info-boxes?
They should absolutely not need to understand everything about Wikidata. This would not work. But for those who interact with Wikidata regularly it should be easy to understand what is going on - not in detail but the bigger picture. To get there we need to improve a few things in the user interface but we will also need to adapt our help pages to be less technical.
What can people do to make Wikidata be useful in their language
Wikidata is inherently a multilingual project. It allows you to use the site in your language and will show you the data it has in your language. Have a look at for example item Q2 in German, English and Spanish. To be able to do this Wikidata needs to know the names of all these things in the different languages. That's what we call "labels". These labels are really important in all kinds of places in Wikidata as they make it possible to refer to things by their name instead of the identifier we gave it - Q2 in the example above. So the best way to make Wikidata more useful in a language is to enter a lot of labels and descriptions in that language. The Special Pages Entities without Label  and Entities without Description are there to help with this as well as the Terminator tool that sorts them by how often they are used to increase effectiveness. This is especially important for the smaller languages. If you speak several languages you should also add a babel box to your user page on Wikidata like I have done on mine. Once Wikidata knows which languages you are speaking this way it will show you the labels and descriptions for a given thing in these languages as well and will let you complete them in case they are still missing. Over time Wikidata will become more and more useful in all languages we support. One of Wikidata's major goal is supporting smaller Wikipedias. This is one of the most important steps on the way to get us there.
People like Magnus and Denny visualised data that exists in Wikidata, how important do you think visualisation is?


Visualisations are crucial for Wikidata. They are a way for us humans to make sense of the vast amount of data. They allow us to see patterns. They allow us to see where we are missing information. They allow us to see where we have outliers in the data that need closer examination. (Belgium has the information that it shares a border with Australia? Probably not when you look at it on a map...) These things are a lot easier to spot and make sense of when you have a nice visual representation of them. But they also show us how far we have already come and give us a sense of achievement. Look at this gif for example. It shows the progress of adding geocoordinates to Wikidata over the first days this was possible. And last but not less important visualisations are of course beautiful and fun.
What can be done to help people be effective in Wikidata
Make it easy for them to get started and understand the basic concepts quickly. Make it easy for them to find like-minded people for example in the task forces. Keep the number of rules low. Create more tools like Terminator and more Database reports to make it easy to find the areas that need more work.
Do you consider that Wikidata is a project in its own right or is it beholden to the Wikipedias, particularly the biggies?
Wikipedia is definitely the most important use-case for a long time to come. However as we're already seeing now Wikidata's data is of use for many many parties outside Wikimedia as well. This will only increase. It is definitely a project in its own right.
What is your dearest memory of Denny as your predecessor?
I have many dear memories. When we met for the first time a few years ago it was in a small room at our university in Karlsruhe to discuss Semantic MediaWiki and its community and development. It struck me that both Denny and Markus (the two founders of Semantic MediaWiki) just got it. They understood what it means to build a community and develop a project in the open with all its benefits and drawbacks. A very rare trait I can tell you. Since then we've built this amazing project and an incredible community gathered around it. Along the way we've been to Wikimania in Washington, D.C. and Hong Kong and many other events. At each of those events we've met incredible people who are passionate about what we're doing and willing to help. I'll never forget that and I'm looking for more of it to come.

Friday, October 04, 2013

#Mobile, a #Wikimedia #investment paying off

To measure if a project executed by the Wikimedia Foundation is paying off, it is not the financials that are to be studied. For the WMF to be more successful, it has to increase its reach and bring more information to more people.

As more people are moving towards "mobile" devices, it is important to be credible on all devices that are categorised as mobile. That takes different technology and all new technology needs to be implemented without disrupting the main business.

The numbers look great; 74% growth YoY. This includes the wonderful growth due to Wikipedia Zero, it is obvious that the quality needed to support mobile devices is increasingly relevant.

There are many components to MediaWiki being developed and there must be a roadmap to bring much of it together. It does include everything from making sure that content in all the languages can be read on a mobile to adding and editing content.

What makes the achievement of this roadmap so hard are all these little DIY projects that paid off so well for its users. They hate to lose their tools. As technology shifts, they invariably will suffer. The main goal however is to get more people to use, interact, edit with all the Wikimedia content. It is hard work, a struggle and it is continually a work in progress.
Thanks,
        GerardM

Thursday, October 03, 2013

Reaching out particularly to all the #USA Secret Service #Analysts

Many employees of the USA government have been send home. They have time on their hand and really relevant expertise.

I would like to reach out to all of them and invite them to edit Wikipedia and Wikidata.

Many of the analysts of the many secret services have been send home as well. I would like them to help us expand the information about all the civilisations they know so well.

I do not ask them to share anything that is secret. I do ask them to consider the need for a neutral point of view and provide sources. I want to remind all my fellow Wikimedians that our coverage about the Arabic and Islamic civilisations is not that great. We need better information and these people have time on their hand.

Please help us bring the sum of all knowledge together so that people can understand each other.
Thanks,
      GerardM

How to connect to #Islam


#UNESCO supports #online access to library content on the Arab, Egyptian and Islamic civilisations. As it is, it contains some 155,000 volumes. Even better, a module that will be available on the Internet is being developed for any researcher who could enrich the catalogue by data of contextualization.

When you consider the information on both Wikipedia and Wikidata about these subjects are not of the same standard, not as informative. Consider for instance the article on the Mahdian Crusade, only the Europeans are mentioned as "notable participants" while elsewhere in the article Islamic notables are mentioned by name.

Similarly subjects and people relevant to the Arabic, Egyptian and Islamic civilisations are not as well presented in external databases like VIAF. It would be cool when libraries like the IDEO library in Cairo or the library of Alexandria were connected as well. If it cannot be in VIAF, how can we enable more of the relevant literature and information they hold in Wikidata?
Thanks,
      GerardM

Tuesday, October 01, 2013

The #disgrace of the USA #Congress


Battle of the sexes, the #Wikidata story

Regularly studies are published indicating how many #Wikipedia articles exist about women compared to the number of articles about men. These studies serve their purpose.

A lot of data is collected to make this possible, the articles about a man or a woman. Obviously every one of those articles is about a person.

The most obvious thing to do next is use the collected date and update Wikidata items where needed. Less obvious is to compare the data again based on what Wikidata has to add.

One of the most statements I have made are "person" and "male"/"female". Many of the items had links to multiple Wikipedias. With more data coming into Wikidata, it is highly likely that the combined data will give better precision.

An added benefit is that it will make such a comparison possible for the "other" Wikipedias, the ones that are not studied individually in this way.
Thanks,
      GerardM

What qualifies as an "Emir of #Morocco"

When you specify things in #Wikidata, some questions are best dodged. For instance when did Morocco become Morocco or England England? Questions like these are particularly relevant when one dynasty replaces another.

When you look at the current border, you can imagine that the border changed rapidly particularly caused by the upheaval of the change of the dynasties. With the change of a dynasty big changes happen. The question is, how does it affect the continuity of a country or its people is not easily expressed in Wikipedia let alone in Wikidata.

Options:
  • both emirs of the Saadi dynasty and emirs of the Alaouite dynasty are "Emirs of Morocco"
    • however they do not succeed from one dynasty to the other on the person
  • the rulers, their dynasty are associated with maps like in this example from an Algerian university
    • to do this, we need dated maps or dated animated maps
  • annotate the "Emir of Morocco" with both dynasties but not with the country
  • you go with the autonym of the country at the time
Whatever is to be done, Wikidata is very much connected to many idiosyncratic Wikipedians who all may bring their strong opinions on such matters to the table. Such discussions are needed as much as dynamic maps are needed to explain the dynamics of power.
Thanks,
      GerardM

Monday, September 30, 2013

#Wikidata qualifiers

When you state that someone is a "Fatimid Caliph", it helps when you can qualify such a statement. The Fatimid Caliphate ended centuries ago so it helps when you can state the start date, the predecessor, the end date and the successor of that Fatimid Caliph.

Qualifiers are particularly relevant when there are multiple "offices held" by a person. It is customary for Islamic nobility to gain experience as a governor of a part of the realm. This ensured relevant experience it also was the start of a power base and lead frequently to insurrections.

The current iteration of the qualifier functionality is powerful. It can do with several improvements.
  • every qualifier deserves its own source. 
  • when standard qualifiers are identified, it makes sense to have them in a particular order
  • the software harvesting information for Wikidata is not yet qualified to enter information properly
  • external data resources have to learn how to deal with Wikidata's qualifiers
At this time the only option is entering data manually. As you read the articles, you learn about so many factoids. You also learn how much work can be done to make Wikidata shine even brighter.
Thanks,
       GerardM

Sunday, September 29, 2013

Red links on the #Arabic #Wikipedia; the sultans of #Sennar

On #Wikidata, you will find a succession of the sultans of Sennar. The information provided is limited because the articles about them are limited. There is no date of birth or death information and for only a few it is known who the fathers, the suns are.

On the Arabic Wikipedia, the article about Sennar has red-links for the Sultans. When you include them on the articles at Wikidata, do they become "infra red"-links?

Sennar is an important part of the history of Sudan, and consequently it is relevant to have information in Arabic. Copying in the red links as labels seems obvious. The question is how to make it more obvious for the Arabic Wikipedians that Wikidata has information relevant to them.
Thanks,
      GerardM

Saturday, September 28, 2013

Amara Dungas, a sultan of #Sennar

Sennar is geographically part of what is now #Sudan. The history of the Sudan seems to be one of conflict with the sultanate of Sennar providing some continuity for over three centuries.

When you read about it, Sennar was a backwater and this allowed it to develop its own culture which was distinct. As so few people are interested in this culture, these people you will not find much information in Wikipedia.

The information you do find of the Sennar sultans in the English Wikipedia is problematic when you want to enter the data in Wikidata. The problem is not only that there is no article for many sultans, the biggest problem is with the dates.



The date when the reign of Amara Dungas ended is given as AH 940. This date is provided in the Arabic calendar and as the illustration indicates, this is either 1533 or 1534 AC.

Technically an acceptable solution is not that hard; when it is possible to say 1533s, meaning it can be up to one year earlier or later it would already be a big improvement. A proper solution means that dates can be entered in the original calendar. This is a lot more complicated but it will allow people to enter more accurate information.
Thanks,
       GerardM

Friday, September 27, 2013

Bye


Importing data from the #Polish #Wikipedia

All new information in #Wikidata has an origin. It can come from many sources and the quality varies. When Matmarex mentioned his source as a Polish project about information about persons, I wanted to learn more.

This is the kind of project that we should welcome at Wikidata. Please have a read and be happy with great undertakings like this.
Thanks,
       GerardM

What is the data you are importing based on
The data is based on the index of biographies maintained by hand by a few dedicated Polish wikipedians at Noty_biograficzne . The nature of created-by-hand data is that not all of it can be automatically parsed, but it was surprisingly consistent – accepting just several common variants resulted in over 60 000 items the bot could understand and only several hundred that could not be used (I am hoping to sort these out by hand). Some typos in the source are unavoidable, but overall the quality seems to be very high.
Getting quality personal information has been a project on the pl.wp for quite some time, can you explain what it means ?
I am not sure myself what the index was intended to be, not yet being a wikipedian when it was started in 2004 – possibly a crossover between a category system (the concept of a category was only introduced on that year, I don't know what was first) and a list of articles needing creation. Currently it serves as, well, an index – list of all biographies on the Polish Wikipedia, ordered alphabetically by last name (or, in some cases, by pseudonym). It's easier to find what you're looking for if you only remember last name of a person and possibly their occupation than using the built-in search system (for example search suggestions are ordered by article title – thus first name – and it's not possible to limit results to only biographies).
You are running a bot adding descriptions in Polish, what software are you using..
Unlike most bot operators I'm not using the Pywikibot framework – I opted for my own custom-written library in Ruby called Sunflower and a a set of scripts using it.
Does it use the Wikidata API and why is this important
Yes, both for uploading and using the information. The API is basically what made the project possible.
What other data do you have about all these people
The old index contains birth and death years in addition to the descriptions. I didn't upload it because it's basically unsourced (and unsourced data is seemingly as frowned-upon on Wikidata as it is on Wikipedia, if not more) and because, when I tried comparing them with birth and death categories on the biographies themselves, I found over 1500 conflicts.
Can you use your data to compare against the data on Wikidata
Not really; there isn't much to compare in this area, especially since I was uploading basically free-form text. Only several hundred items out of the 60 thousand my bot edited already had a description in Polish, me and a couple other editors reviewed them all in a few days.
The birth/death data could be compared, though, but I haven't looked into it. Any help would be welcome!
Can you add your data where Wikidata has none
There are a few things the index uses that are not yet present on Wikidata. The birth and death dates are the biggest one, real names of people using pseudonyms (such as Sting or Madonna) would be a valuable piece of information as well. I didn't try to upload either – the dates would need better sources (the 1500 conflicts are a strong indicator that information from Wikipedia might not be good enough) and there are currently no properties defined for first / last names because of how complicated the topic is (there is currently a discussion under way).
Did you know that this type of data becomes available on several Wikipedias in stub articles
I considered parsing the articles themselves to extract the descriptions, but decided that this would be too error-prone to automate entirely. Instead I developed a gadget that helps users write short descriptions for biographic articles by extracting the information from the lead-in paragraph and presenting it on the index pages – they can be adjusted by a human and saved to Wikidata with one click! This benefits both projects at once and I think is a good example of how they can work together.
Is it possible to transfer this biography project from the pl.wp to Wikidata
It could be done entirely on Wikidata and using Wikidata information, but there are two preconditions – presence of the required information on Wikidata (birth/death dates, last names for correct sorting) and ability to generate lists from the data (so-called "phase 3" could accomplish this).

Thursday, September 26, 2013

Thank you Denny

Denny has been the main man at #Wikidata. He has done a splendid job. I have known Denny for many years and I admire his achievements. One fond memory was meeting in Rome because we both happened to be there :)

Denny will join this big search company and I am sure we will have a friend in Denny at Google.
Thanks you,
      Gerard

#MathJax is localised at #Translatewiki

#Wikipedia has potentially articles on any subject. When it is about other countries, MediaWiki supports fonts when needed. When it includes mathematics, there is another need; the need to display formula properly. Mathjax has as its slogan: "beautiful math in all browsers" and it does do a great job.

The one thing that was holding back adoption of MathJax was its lack of support of other languages.

Now that the community of localisers at translatewiki.net has adopted MathJax it is a different story. The lack of a localisation can be fixed in the time honoured way; you can do it. There is not only support for the localisation in potentially 300+ languages, there is also strong internationalisation support available when needed.

When your Open Source software needs an international audience, consider translatewiki.
Thanks,
      GerardM