Monday, August 12, 2013

What #Commons can learn from #Shutterstock


When you are in need for an #illustration you can find it at Shutterstock. When you are willing to contribute illustration to a project like Wikipedia or Wikivoyage, you can do that at Commons.

At Shutterstock they have over 25 million images and at Commons there are almost 18 million images. With such big numbers finding the perfect image is a serious problem. Shutterstock is optimised for finding images; it pays their rent.

Arguably, Commons has a much better coverage of subjects both in time and in geography. Try Kiribati for instance..


The reason why Shutterstock is worth billions of dollars is because people download images and use them for their purpose. As far as I am aware the Wikimedia Foundation does not know how often an image is downloaded from Commons.

There are so many reasons why Commons could be much more relevant. Relevance comes with use and people will use Commons when they can find their perfect picture.
Thanks,
      GerardM

Sunday, August 11, 2013

#Moroco is a modern country

When you look at history from the present, it is normal to look for continuity. The history books are written today and they inform about history from a modern perspective. 

When you read an article like "List of rulers of Morocco" the list includes the historical precursors to the modern state. These precursors include to some extend the same territory but the make up of the people can be really different. 

The Idrisid, the Almoravid, the Almohad, the Maranid and the Wattasid dynasties were Berber. Many of them were Shia muslims. Contrast that with the later Arab dynasties and a country that is predominantly Sunni.

It is relevant because it does not make sense to state that an Idrisid ruler is a "Morocon head of state". The question is very much how did the Idrisids or any of the other dynasties call their country because that is what they were ruling.


Another thing to consider is what happens when a country is conquered. It does not make sense to suggest that there is some kind of succession. The reality is that both the dynasty and the country become irrelevant as they are absorbed by the conquering nation.

What is obvious to me is that many of the solutions as used in Wikipedia are based on template logic. With Wikidata we have the opportunity to come up with better solutions. What they will be is not immediately obvious.
Thanks,
       GerardM

Saturday, August 10, 2013

Using qualifiers in #Wikidata

I have been adding statements to many Wikidata items. Lately I have been adding information about several dynasties. So far I have added statements like "preceded by" and "succeeded by" and expected that people would appreciate in what way people were preceded or succeeded. I did not indicate when this happened.


Doing it this way did not really satisfy me. I found a great example of how this could be done and I copied it. In my opinion what you can see above is a big improvement.

However..

When you indicate that a person is the "head of state" of a country, it makes sense to have the same information available on the country as well. Purists may object to the fact that this information will be stored twice. I have no idea how to solve that.
Thanks,
      GerardM

Tuesday, August 06, 2013

What to do when #Wikipedia does not provide the information

King Abd al-Rahman of #Morocco was a nephew of the king that preceded him. His father of mother were probably brothers of king Slimane. As long as I do not know exactly how they are related, the genealogy tool made by Magnus will not show all the rulers of the Alaouite dynasty.

Wikipedia does not have this information.
  • I googled
  • I went to the library
  • I asked on Facebook in a SIG with many people from Morocco
So far I do not have my answer. I am really curious what the best way is for finding answers to a question like this one.
Thanks,
        GerardM

Sunday, August 04, 2013

The case against the use of the #GND classification

#Wikidata uses the GND classification and twice it has been proposed to get rid of its use. So far the issue has been unresolved.

To understand the GND classification, you have to appreciate its origin. This classification has been created and is maintained by the Deutsche National Bibliothek. The consequence is that we cannot modify this classification to suit our needs and still call it the "GND classification.

When a library creates a classification for its own purposes, it is perfectly all right to classify "typhus" as a "Term". From a library point of view it is and publications may refer to it. When you look at it from the use of something like Wikipedia, it is likely to be categorised as something like a bacterial disease.

In this context, we can assume that all bacterial diseases are caused by a bacterium. This bacteria has a name. The disease will be known in one or more medical classifications. It being a disease, it will have all kinds of other attributes and none of them are easily associated with the classification as a "term".

The one thing we should retain of the GND classification in Wikidata is the identifier of a Wikidata item as used by the GND classification. This will allow us to compare information with the Deutsche National Bibliothek.

Obviously, once it has been decided to ditch the GND classification in Wikidata, we need to carefully migrate to a classification that is more appropriate.
Thanks,
         GerardM

The battle of Verdun in #Wikidata

In a previous blog post I wrote about dates and my wish for entering and displaying dates in the Islamic calendar. As a result there were several interesting reactions.

For some people it was news that Wikidata supports dates. The question put to me was: "This Q130847 has geo coordinates but no date". As this refers to the battle of Verdun, I replied by adding the start and end date to the Wikidata item for the battle of Verdun.

One bit of good news is that MediaWiki does support the Islamic calendar, as a consequence the issue of using the calendar got a new dimension. I was also asked if other calendars like the Balinese calendar could be supported. In principle MediaWiki should support any calendar however, I would prefer to start with the support of calendars that MediaWiki supports in Wikidata.

The bug for support for the Islamic calendar in Wikidata was given the I18N label in Bugzilla. I hope that our language team will take an interest in supporting calendars in Wikidata

As it is fairly obvious that the battle of Verdun is a battle, I added "Instance of" "battle". When you look where battle links in Wikidata you will find many more battles that are linked to the concept of battle. I added a few more battles that I just happened to know.

When you look at the info-box for the battle of Verdun, you find many bits of information that we should be able to support in Wikidata. It would be cool when we can.
Thanks,
       GerardM

Friday, August 02, 2013

#Ramadan 1434

Not that long ago #Wikidata gained the functionality of adding dates as an assertion to an item. It is really nice because it is now possible to indicate when a person was born or died.

At this moment we are in the last week of the month of Ramadan, in the year 1434. When you check a calendar forthis month, you may notice that Ramadan is partly in the month of July and August 2013.

When you are adding assertions to Wikidata relating to Islam or Arabic history, you will find that many sources use the Islamic calendar and not the Gregorian dates. As a consequence you have to convert them to Gregorian before you can enter them into Wikidata.

To solve this Wikidata could accept Islamic dates. It just takes some programming. With a little bit of help from Google you can find PHP code like I did. It will probably provide basic functionality. I cannot tell.

The question is, is there someone who is willing to write the code in Wikidata so that it will provide support for Islamic dates? I am sure there will be all kind of issues to resolve.
Thanks,
       GerardM

Use #Wikidata in the #UI of #Wikipedia


Some people do not like the Visual Editor. They think it is not mature enough. It has been explained to them how important it is for the developers to get continuous feedback. The perfect VE will not come fully mature out of the head of Zeus like Pallas Athena did. It is more likely that it will take continuous incremental improvements from being a product with great promise to a great product.

One of the arguments used against the VE is that it will not be possible to include a button or something for all the templates that are in use on a Wikipedia.

This is true at this time. However, it is possible to narrow this number down substantially. This is done by looking at the context of an article. To the right you find the current genealogical map of the Ottoman dynasty. For all these people many things are known in Wikidata and for many of these people one or more templates are already in use.

When we know that a specific article is about a certain type of topic, we can exclude most templates because they are not relevant. Conversely when we know that a specific template is in use on a Wikipedia, we can suggest the use of this template in another Wikipedia.

My point is not so much that this can be done but more that with the Visual Editor new opportunities become available. Such opportunities are not what people are used to but they deserve consideration. They will enable more people to improve their Wikipedia particularly in all the "other" languages.
Thanks,
       GerardM

Wednesday, July 24, 2013

#Wikidata IS a multilingual project

When information is not available in English, the genealogy took falls back to its identifier. All it takes to remedy the situation is to add value by inserting an English label.

In this instance it is all about the Ottoman dynasty and as you can see, Turkish spelling is used in English. The Turkish link for Q6355649 goes to Hümaşah Sultan. That will do as an English label as well.

The one thing I would dearly like is to be able to use the genealogy tool in multiple languages. This is how people can check if labels exist in another language.
Thanks,
       GerardM

Tuesday, July 02, 2013

Writing #Cuneiform on #Wikipedia

When you always wanted to write in the cuneiform script on Wikipedia, now you can. Technically there is not that much to it.

{{Cuneiform|{{linktext|𒄖|𒉈|𒅁|𒌨|𒅎}}}}

This is what it took to get the text you see on the infobox of the cuneiform article.
....
This is fun, nice but not all there is to it. What is really exciting is that you can now write in many, many scripts and expect that the person will see the text properly.

It is no longer necessary to create a screenshot. If anything, this is the time to find all those screenshots and replace them with the proper text and invoke the ULS and its webfonts.
Thanks,
      GerardM

#OpenDyslexic provides support on the English #Wikipedia


Today the "Universal Language Selector" premièred on the English Wikipedia. There is a ton of functionality in there and it has a lot of potential. The one thing that may prove to be a game changer for people with dyslexia is the inclusion of the OpenDyslexic font.

Once people with dyslexia start to adopt this font, chances are that they can actually read/use Wikipedia. A lot of people are dyslexic; to quote the en.wp article on the subject: "It is believed the prevalence of dyslexia is around 5-10 percent of a given population although there have been no studies to indicate an accurate percentage".

The one thing left to do; let people know that Wikipedia has improved usability for people with dyslexia. Please tell it to the teachers, the parents and most importantly to all the people you know who are dyslexic.
Thanks,
      GerardM

NB it is also available for many other languages on many other Wikis.

Monday, June 24, 2013

Homo sapiens anyone?

There is a big thing about bot generated articles based on the taxonomy of species. My take is that this whole discussion is of the rails. It is off the rails because people do NOT understand the vagaries of taxonomy. It is off the rails because people forget what we aim to do; providing knowledge by sharing information.

I loved what Erik Zachte had to say; he wrote about uploading images to Commons and finding that there are no articles in any Wikipedia about the subject of the pictures. It changed his perspective on bot generated articles. When enough information is available to a bot, it can generate articles on 280+ Wikipedias. The alternative is not providing information on the subject in any language.

Another loud argument was about taxonomy; article nbr 1,000,000 on the Swedish Wikipedia is about a species that was recently renamed. As some people would have it, the information was no longer “valid”. One counter argument is, when people know a specimen by the “old name” there would be no information to be had. Another counter argument: from a taxonomy point of view validity of a name is only in the quality of the publication and as a consequence, the old name is valid. To make this point abundantly clear; Homo sapiens is what most people know for the taxonomical name for a human being. I am not completely sure and I care not that much but I seem to remember that “Homo sapiens sapiens” is what has been used more recently in taxonomy for us "thinking men".

Let’s cut the crap and analyse the situation:
  • Many Wikipedians hate stubs without any consideration for the opinion of others
  • Stubs, particularly well-designed stubs are an invitation to edit them
  •  Our prime objective is to provide information
  • In all the recent huha there has been little talk about technical possibilities

One solution for machine generated stubs is to have them in their own namespace and move them with the first human generated edit. This will not shut up all the detractors but it removes their arguments from them.

Another solution is to have the bots generate the information only when requested. It does not need to be saved, it only needs to be cached. Given that it is a bot generating the information, the script it uses can be translated for use in other languages as well.

Yet another solution is to have such scripts associated with Wikidata items. The information provided in this way would be truly complementary to what is available in Wikipedia. An added bonus would be that it will take away any room for Wikipedians to complain. Hm.. possibly, probably not.
Thanks,
      Gerard

Saturday, June 15, 2013

Red links in #Wikipedia categories are possible… use #Wikidata

I have been looking for lists on a given subject. My requirements are “simple”;
  • I want a complete list
  • I want to know if there is an article in a language on a Wikipedia for the items on the list
The first requirement excludes categories.

The second requirement is provided to some extend by categories. In this case I used information from the English language Wikipedia. There was an article on the subject as well and it contained a list and this list contained items that were not in the category. I read many of the articles to add statements on these persons and based on the articles I had to conclude that some of the category items and list items were wrong.

All this is to be expected; Wikipedia does not claim to be 100% correct. It is for the people working on the content to refine and improve the content. The funny thing is that by reading articles I found candidates for other “categories”.

To get a more complete list, you can iterate the process on other Wikipedias. In addition you can add items to Wikidata based on “external” information. It takes some effort and, in a perfect world you would be aware of any items that are already known in Wikidata.

A list that is compiled in this way is superior to what categories offer; you would have red links in many places but you can provide information using info-boxes using the Wikidata statements. 
Thanks,
      GerardM

Sunday, June 09, 2013

Thanks for all the fish

Milosh is leaving the movement. It takes me aback. Last time I met him was at the Amsterdam hackathon, we discussed several things that might help our shared dream of more linguistic diversity in the Wikimedia movement.
  • Wikidata is to be opened to any and all languages that are recognised as a language. This includes artificial languages and dead languages. Some people may find it controversial that we consider Wikidata to be a project in its own right and, the connection to Wikipedia incidental.
  • Wikisource should be one project like Commons and Wikidata. It would be best when all the instances of Wikisource are merged. Wikisource as we know it, is very much a workbench. Important is that people know how to use the tools. It makes no difference what language your user interface has, what matters is the language of the text. This can be any language
  • When the work is "done" on Wikisource, it should be advertised, it should be marketed, it should find a public. That is very much something best done OUTSIDE of Wikisource. 
Ah well, these things make sense, they will improve linguistic diversity and they will create an environment that will even stimulate the creation of more Wikipedias. The question is how to find the energy to make this happen.. The Dutch are known for their windmills, just like the Spanish. We will miss Milosh for this fight.
Thanks,
      GerardM

Tuesday, June 04, 2013

#OmegaWiki supports #OpenDyslexic

#OmegaWiki aims to be useful as a platform. Much of its functionality is in the extension that allows us to be special. The core functionality however is provided by MediaWiki. The Universal Language Selector is an extension of MediaWiki waiting in the wings to become available on all the wikis of the Wikimedia Foundation.


Part of the ULS is the use of webfonts and, OpenDyslexic is a font available for many languages that use the Latin script. Languages like English, French, German, Dutch... The font looks odd to most people but for many people who are dyslexic it makes text actually .. readable.
Thanks,
     GerardM

Monday, June 03, 2013

#Defending the Wiki in #Wikidata

When I was younger, #Wikipedia was the encyclopaedia everyone could edit. I could because it was a wiki. I could start new articles, make changes to articles and be happily productive in this way. Sure, sometimes I made mistakes but I was not the only one around and they were as happily productive as I was fixing things after me and others.

With the call for citations things became more formal but it is still possible to some extend to just write and edit articles.

Wikidata is about assertions. I love the definition of assertion: "A positive statement or declaration, often without support or reason". I love the definition because it is what allows Wikidata to be a Wiki. When an assertion can at first be added in Wikidata without support or reason, people can add assertions freely. Certainly when you assume good faith you will be happy when people do exactly this.

When people add information using bots, they have a source that provides them with the assertions. Quite often these assertions originate in one of the Wikipedias. Alternatively they come from sources that are happy to share their information. While I applaud the sharing of data between sources, I believe quite strongly that importing the structure and limitations of external sources will hamper the development of Wikidata.

I routinely add "main type (GND)" with the value "person" when an item in Wikidata is about a person. However, the other values that are associated with this "main type (GND)" are absolutely horrible to the extend of unusable. Adding a link to the GND database is how Wikidata can add value to its usefulness.

As Wikidata is a wiki, I use the attributes available to the extend they make sense to me. Many attributes are lacking and the procedures for getting attributes are not exactly easy or obvious. (I did request an additional attribute called "rada"). The point is that these lengthy procedures make Wikidata less of a Wiki.

Wikidata is a Wiki and consequently people are free to add "statements". Adding a requirement of sources to any and all assertions are absolutely counter productive because we can only improve assertions once they have been made. External requirements like this will effectively kill a Wikidata community. It will also ensure that the Wiki part of Wikidata is a lie.
Thanks,
       GerardM

Thursday, May 30, 2013

#WMhack - produce a list that shares a #Wikidata attribute

At the Amsterdam hackathon it became clear that Wikidata can be used as a powerful tool to improve Wikipedia. The idea of a hackathon is that you go home with new hacks and, I want to share a really nice one. It is a list. A list that indicates every known link within Wikidata from a given topic.

The cool thing is that it is the standard MediaWiki "what links here" functionality.

When you copy the "following pages" for this item, you can use the "text to columns" functionality to get something that you can use. With search and replace you can build a string that looks like this..
 {{#invoke:Available|link|Q48210}}
Many of such strings together produce a list when the Lua module Available is present on your Wiki.

When you see the Q numbers, there is no article about the subject. When you see a name in red, the label exists in Wikidata but there is no article registered for this subject. When you see a link in blue, there is a label and, there is an article.

This is a genuine hack that is really useful to complete articles on the same set of Wikidata items in any and all Wikipedias.
Thanks,
       GerardM

#WMhack - #SignWriting in #Wikimedia #Incubator

One visitor at the Amsterdam hackathon was Stephen E Slevinski Jr. He was a man with a mission. His mission was to get American Sign Language on the Incubator. As you can see in the picture below, he succeeded in his mission.

Several members of the language committee knew that it was possible for Stephen to succeed. We also knew that when this hurdle was taken, any and all sign languages can ask for their own Wikipedia. There have been requests for languages like Danish sign language in the past as well.

Now that we have the technical ability to show a sign language, the old reason not to express "eligibility" for these languages has gone. Any and all sign languages that have an ISO-639-3 code will find that there is nothing that will prevent them to work towards a Wikipedia or any other project in their language.


I want to thank a few people. First all the people at the Amsterdam Hackathon who helped Stephen make this dream come true. Secondly, I want to thank Stephen, Valerie and all the other people at the SignWriting Foundation for their continued belief that their languages and their culture deserve their own  Wikipedias.
Thanks,
       GerardM

Sunday, May 26, 2013

#WMhack - Speedy deletion

Ah well, #Wikipedia and its speedy deletion process... according to its rules " This criterion does not apply to pages in the user namespace, nor does it apply to valid but unused or duplicate templates". As you can see, this page is obviously part of my user page.

Wikidata has information on a specific group of articles that may or may not exist on this Wikipedia.


PS The issue has been resolved; the page is no longer in line of immediate deletion .... pffff

#WMhack - Dates for #Wikidata

At a hackathon you may preview what is about to happen. Dates for Wikidata is something that has been eagerly awaited. Given that Wikipedia has articles on many events, it is obvious that Wikidata is one place where assertions on these events should end up.

The most obvious dates are the date of birth or death for a person or the date a ruler started his rule end, the date it ended.

There are many issues that I can think off about how dates are to be used. But the fact that dates are finally being tested and are likely to arrive in a weeks time is a reason to be happy.
Thanks,
      GerardM