Thursday, December 29, 2016

#Wikidata - Khagan of the Rouran


A great sign that Wikidata gains traction in other languages; much of the data for Yujiulü Shelun, Khagan of the Rouran from 402 to 410, does not have labels in English. When the idea is to include all these Khagans of the Rouran it becomes a challenge. The English article does have many names but do they fit what is already there for other languages.

The challenge is to do good and bring things together. It is relevant to have all the right items properly connected. One thing that is missing; the item for Khagan of the Rouran. That is easily fixed.
Thanks,
      GerardM

Tuesday, December 27, 2016

#Wikimedia and the "official point of view"

One of the pillars of #Wikipedia is its Neutral Point of View (NPOV). The point is that we should not take sides in an argument but should present arguments from both ends and thereby remain neutral. The problem is what to do when arguments are manifestly wrong. When science repeatedly shows that there is no merit in a point of view.

What to do when it is even worse, when science is manipulated to show what is of benefit to some. When the Wikimedia Foundation had its collaboration with Cochrane, it was onto something important. Cochrane is big on debunking bad science.

The new government of the USA has a reputation that precedes its actions. It already states that science is bad. It will state its point of view. They will argue that it is good for all but how will they substantiate this? In the mean time much of what science said so far will remain standing. The snake oil salesmen will try to sell you their product and I wonder how it will find its way in Wikipedia. Will we look at science and will we resist the snake oil?
Thanks,
      GerardM

Monday, December 26, 2016

#Wikidata - #caste and how to include it in Wikidata

With all respect to cultural heritage, forcing people to be included in any caste is a form of discrimination. In an article about Nangeli, a woman of the Nadars, it becomes clear how important it is to understand its history

The Nadars are a heterogeneous group, comprising people of diverse standing. When in a school curriculum the story of Nangeli was included, it did not do justice to this diversity.

The problem with discrimination is that it has to be simple or it is not understood. It is how I interpret why it was pulled from the curriculum. This whole notion of the impossibility of there being one simple caste system is expressed well in the Wikipedia article on a historic article on the Nadars; the Sivakasi riots: "This belief, that the Nadars had been the kings of Tamil Nadu, became the dogma of the Nadar community in the 19th century". It casts doubt on schema where castes are expressed in a simple way.

What we can do is linking what we know is related. Link historic facts associated with class and castes. But it starts with making the effort.
Thanks,
      GerardM

Sunday, December 25, 2016

#Wikidata - the grandson of King Thibaw

When the BBC writes an article about royalty, it makes sense for both Wikipedia and Wikidata to have correct information available.

Descendants of king Thibaw Min it helps when it is known that this king was part of the Konbaung Dynasty and that a dynasty is not a country.  This is relevant because any claim to Myanmar is based on being part of that dynasty.

It is simple; dynasty is family. It is why Mr Trump and his offspring are factually a business dynasty.. When we are to get our facts straight, it makes sense to understand such basics. A dynasty can lose control over its "assets" but it remains a family.

Historically there have been many families with claims to a crown. Understanding such a claim is of interest and it is relevant to know the history of the whole world. History is not only lived in the western world.'
Thanks,
      GerardM

Tuesday, December 20, 2016

#Wikidata - a country is not a dynasty

When a "country" comes into being, it is after a struggle. In the same way when a "country" comes to an end, it is after a struggle. The same is true for dynasties; when a royal line comes to a start or an end, it is not without a struggle. However sometimes in a country there is continuation and one dynasty follows a previous one. Several dynasties succeeded each other in the Delhi Sultanate. The "country" finally ended with the last of the Lodi dynasty.

So when a country knows only one dynasty and starts and ends with that dynasty, it does not make the dynasty the country. Making up a name for a country is easy; when these monarchs are called "king" it is a kingdom, when they are a "sultan", it is a sultanate.

If there is one drawback, it is that there might be a name for that country in the languages of the people who were linked to it. For this reason all the countries that I am about to create may be prime suspects for a merger.. The item, not the country :)
Thanks,
     GerardM

Saturday, December 10, 2016

#Wikidata - Sembiyan Mahadevi - is it a title or is she a queen?

Queen Sembiyan Mahadevi was the spouse of  Gandaraditya, her son was Uttama Chola. Many of the Chola queens who followed her used "Sembiyan Mahadevi" as a title. This is what the English article tells us.

To really accept that it was a title, a source would help. It would be cool to have a list of all the people who used the title and it would be good to separate the person from the title in separate articles. It seems that the Tamil article is more substantial but as I do not read Tamil and Google translate does not help me sufficiently to understand what it says. 

Queen Sembiyan Mahadevi matters not only because she is important in the Chola dynasty but also because of the relevance she has in Tamil culture. Her father was a Mazhavarayar chieftain but Wikipedia does not know about them. 

When Wikidata knows about Indian nobility, its dates and connections, it becomes a resource that is helpful. Once her father has a name and it is clear what is meant by a "Mazhavarayar chieftain", slowly but surely it becomes clear who ruled where and who were contemporaries. It would be cool when Wikidata allows for a query that shows a "monarch" and shows fellow monarchs in neighbouring countries. 
Thanks,
      GerardM

Thursday, December 08, 2016

Was Cezhiyan Cendana a Pandyan king?

There is no way for me to find out if Cezhiyan Cendan was a Pandyan king or not. The only source I can find is a blog saying so. The problem is that texts in Wikipedia make me doubt. The text in the article for Maravarman Avani Culamani states that he is succeeded by his son Jayantavarman.

One fun fact is that templates do not have sources. It is however what I base information on when I add information to Wikidata. The other interesting point is that dates given are overlapping to the extent that they are not reliable.

So this is where we get into a problem. When information is good enough for a Wikipedia, is it good enough for Wikidata. More importantly is the question how do we curate information like this in a way that helps us all?
Thanks,
     GerardM

Wednesday, December 07, 2016

A Pandya King did not rule #India

The Pandyan Kingdom existed for some fourteen centuries; for many of the kings not much is known; A template contains much of what is known about them; not much.

Arguably; having this information in Wikidata serves a purpose. The information can be curated by people who know about the Pandyan kings and there are several things that they could do.
  • Some of the names of kings seem to be incorrect, certainly inconsistent.
  • The names of these kings can be added in the original language
  • Dates may be added to the period these kings were king
  • The data can be used in one of the other Wikipedias that are relevant in India.
One funny fact is that for all these kings it is impossible to have been a citizen of India. They were citizens of the Panyan kingdom. Many of such facts were added by bot and, it reflects factoids that exist in Wikipedias. It is just wrong.
Thanks,
      GerardM

Tuesday, December 06, 2016

#Research to help #Wikipedia do better

It is one thing to bemoan everything that is problematic with research, it is another to do better. For research on Wikipedia to be published, it has to be about "English" OR it has to be linked to English OR publication is not the end goal.

At the Dutch Wikimedia Conference Professor de Rijke gave the keynote speech. He spoke about the kind of research he is into and he spoke about "Wikipedia" research performed at the University of Amsterdam. He challenged his audience to cooperate and his challenge resulted in me formulating ten proposals for research. The point of these proposals is that I hope they do provide more worthwhile insight and includes a link to “English” in order for it to be published.
  1. Previous research, studied how long it took for a subject to appear in English Wikipedia after it was first mentioned in the news / social media. The new question would be: how long does it take for the same subject to appear in any Wikipedia and, how long does it take and to what extend does it happen for those articles to get corresponding articles in other Wikipedias and how long does it take for the English Wikipedia to take notice?
  2. In the search engine for Wikidata we use the description to help differentiate between homonyms. There are two approaches to a description; many existing descriptions are not helpful and hardly any items have texts exist in all of the 280 languages. There are however automatically generated descriptions. The question is: what do people like more, the automated descriptions or the existing questions? Is there a real difference for people who use Wikidata in English as well?
  3. Many people know their languages, this is obviously true for readers of Wikipedia. For the regulars there is a “Babel” template that allows them to indicate what languages they know. For the others for some purposes geo-location is used to make a guess. Do people find it useful to have it indicated that articles exist in the languages they know in search requests? Does it make a difference that a quality indicator is set for those other texts on the same subject?
  4. Many people make spelling errors when they search for a subject or when they create a wiki link to another subject. Google famously suggests what people may be looking for. We can expand the search and include items from Wikidata (40% increase in reach) but we can also use Google or any other search engine to help people get to the sum of all knowledge. We can ask people to answer some questions after they are done. Are people willing to do this and how does it expand our range of subjects that we know about. Are people willing to curate this information so that we can expand Wikidata and at least recognise the subjects we have no articles about?
  5. When we show the traffic for the articles people edited on in the last month, we gain an insight in what people actually read. We also congratulate people on the work they did and show appreciation. Does this kind of stimulus stimulate more articles? How do you stimulate for subjects that people hardly read (eg Indian nobility).. Do you compare with existing articles in the same category?
  6. There have been several Wikipedias that include bot generated texts. It is a famously divisive issue in the Wikipedia community. There has been no research done on this. With Wikidata there is an alternative way to exploit the underlying data. When the data is included in Wikidata, it is possible to generate text on the fly. This data may be cached for performance issues but there are two main advantages; both the script and the data can be updated. The question is: does it serve a purpose for our readers? Will editors update the data or the script to improve results or will they use the text as a template for new articles? Will it take the heat of the argument of generated texts? How will it affect projects that were not part of the existing controversy and does it work for them?
  7. Wikidata does not allow for the dating of its labels. It follows that it is not easily understood what the relation is between Jakarta and Batavia. How are such issues generally stored as data and what alternatives exist for Wikidata. How does it improve the usefulness of Wikidata as a general topic resource?
  8. Wikidata now includes data from sources like Swiss-Prot. What are the benefits to both parties? Does it make for people editing this data at Wikidata and what is the quality of such edits? Does it get noticed by Swiss Prot and is there a cooperation happening? How is this organised and to what extend does “the community” interfere with the notions of academia? Do such communications exist or are these groups doing “their own thing”?
  9. What is the effect on the ultra small Wikipedias when generated texts are available based on available labels.. Does it mean more interest in creating the templates for articles and work on labelling? What does it mean when such generated articles are available to search engines?
  10. At this time many articles in the English Wikipedia are written by students, university students. The result is positive on many levels but the question is, is what they write understood by Wikipedia readers? When students write their articles, it is mostly based on literature. It is well known that the bias in scientific papers is huge. Negative results are not published and many results from studies are ignored. The question would be: is sufficient weight given to debunking studies or are they put aside with an argument of a “neutral point of view”. This would make sense when students are graded on what they write given accepted fact on the university.

Saturday, November 26, 2016

The problem with #science explained with #Wikipedia

It is a recurring theme. People study a subject and reality is different. The science is flawless, the results are impressive and indeed important strides are made forward. The study of heart disease is a great example; many studies resulted in an improved life expectancy for men. Particularly white men. The Dutch Hartstichting is raising funds for new research because of this existing bias in research. For women in the Netherlands, heart disease is the number one killer because heart disease is different in women; it was not noticed before because heart disease in women was not studied.

Wikipedia as it is commonly known in research has the same problem. It is not Wikipedia as we know it, it is English Wikipedia. My contributions to Wikipedia have not been to English Wikipedia; they went to the Dutch Wikipedia and I will not be noticed as one of the most prolific contributors to Wikimedia projects because my contributions to "Wikipedia" are hardly significant..

As I blogged before; scientific papers do not publish when it does not involve English Wikipedia. The consequence is that when people quote research, their quotes include this bias and strictly speaking it is not necessarily true when you consider Wikipedia. The problem with biased research is that the policies of the WMF are based on the known "facts".

Nothing new so far. We all know it when we are honest. So what can we do to remove some of the bias? The first thing is to devalue any and all research that is English Wikipedia only. It only covers less than half of what we do.The second thing is to evaluate research for its algorithms. When both the algorithms and the data are available, it is possible to run the algorithm on a more inclusive data set and check the validity. With the quality of Wikidata data as a source on all the Wikipedias improving, such an approach is increasingly feasible. The last thing is for the Wikimedia Foundation itself to address this bias, With English Wikipedia being less than 50% of its traffic and workflow, it would be good when a similar percentage of its efforts is focused on the bigger half of what we all do.

So what is the harm? We expect all Wikipedians largely to do what "Wikipedians" do. However, we are not all English Wikipedians. The need other people have is not discussed, not taken seriously. We have seen wonderful examples of potential functionality showcased but it is not taken further, not taken in production because it does not fit the preconceived ideas of what we do, it is not part of the road map. The projects in Wikidata are not about Wikidata but about how to make us all in one big data glob and USING the data is only seen in relation to Wikipedia articles. We do not know how much Wikidata is used, some studies are done but they are in relation to "Wikipedia" and that is not relevant to me. We find that Wikisource gains more and more content that may be valuable to our readers but we do not market this data because we never did marketing for Wikipedia. There are several websites that only do this in a way that could be much improved if we took Wikisource seriously.

It hurts us to only consider English Wikipedia and this bias in research and policy is more damaging than the bias that is considered by the English Wikipedians.
Thanks,
       GerardM

Wednesday, November 23, 2016

#Bias in #research

Actually, it starts with something else. You need to publish so you have to select a subject to study that will be of interest to the publisher..

As a consequence hardly any research is done about the other Wikipedias. I have been informed by a reliable source that it has to be English or it will not be published.

Now Wikimedia Foundation, how about that? Is there any research done on Wikipedia or is all the research biased in this way?
Thanks,
     GerardM

Tuesday, November 01, 2016

#Wikidata year 4; What Gupta year is that?

Wikidata is celebrating its fourth birthday. It is celebrated by some mighty fine gifts. It is a time to reflect on what has gone before and what is ahead of us. Obviously there are challenges we face and my gift are some queries / questions I do not know how to address. I focus on the Gupta empire because it currently has my interest.

During the era of the Gupta empire there was a "Gupta year". An article refers to it and my first question is: what date would the birthdate of Wikidata be in Gupta years?

Obviously there are many maps including the Gupta empire, Can I have them sorted by date please? What other countries border the Gupta empire? Who were its rulers and how does the map change over time?

To get answers is nice but for me it is important that the algorithms involved are relevant to any country old and new. Relevant to timelines old and new. When we can express dates in the "Year Gupta", we can check if dates in Wikidata are indeed Julian or maybe Gregorian..

When we have continuance in maps over time, we will know if a location, a city for instance or the land of a tribe is part of what country; what culture.

Wikidata live long and prosper :)
Thanks,
      GerardM


Saturday, October 29, 2016

#Wikidata - Queen Kumaradevi

Queen Kumaradevi was married to Chandragupta I. According to Wikipedia she was of the Licchavi clan. The coin shows her with her husband on a coin minted by their son.

When you read Wikipedia, you will read about daughters of kings married off to nobility. They paint a picture of alliances, their marriages often meant some stability in an often brutal world.

When you are interested in such things, western nobility is well documented. Not so for nobility of India. I have added lately a series of maharajahs, kings and emperors and am every time amazed that nobody beat me to it. I often document who was related to who and often find missing links documented and add items for them. Regularly the missing links are implied but miss a generation.

I am sure of one thing; India has its fair share of people who know and care about such things. How do we get them interested, how do we get proper information about all this in Wikidata?
Thanks,
     GerardM

Sunday, October 23, 2016

Kigeli V, Mwami of Rwanda

Kigeli was the last ruling Mwami of Rwanda. He died October 16.

When a last ruler dies, it follows that there are previous rulers and, there is a lot that is of interest in the history of the mwamis. His father for instance was deposed because he refused to become catholic.

I have added the rule of several mwamis to Wikidata because such basic information is often lacking. Wikipedia articles are often stubs at best and sources are often absent.


Typically a monarch is part of a dynasty. With a new dynasty it represents often a new family but certainly a change that makes for it to be recognised as such. The article on the kingdom of Rwanda describes the role of the mothers of a king. They are yet unknown to us and consequently a lot of relevant information is missing.

When you see all those red links, it is obvious that significant red links exist in any language. When they are linked to Wikidata, information like the follow up as ruler and who is related to who becomes a task that can be done once and be done well. It is one way to emancipate information that has been of little concern to Wikipedias.
Thanks,
      GerardM


Saturday, October 22, 2016

#Wikidata - statements are doing fine

In September there are more Wikidata items with 10 or more statements than items with no statements. Wikidata is growing up.
Thanks,
     GerardM

Thursday, September 29, 2016

Trust

I read an article, I found what was written astounding and signalled that I had to read it again to really understand what is said and what it implies. The article was published in a quality newspaper; the Independent. The reply that I got was: "Indeed. And it's Fisk, so you can't just pretend it is an obscure journalist talking about something that may have happened..."

As I did not know Robert Fisk, I looked him up. I checked his Wikipedia article and found that he has indeed a reputation that is really good. He received many more rewards than was known at Wikidata so I added several and it is fun to establish the quality of its sources. For the Lannan Cultural Freedom Prize the Lannan website says it all. It is linked on the item for the award and that should suffice. For the Amnesty International UK Media Award it is not so obvious. It is conferred by te UK branch of Amnesty International and it has no dedicated page for the award. I added the award, the chapter and had a look at the pages for the award ceremony for each year. These Wikipedia articles refer to webpages that no longer exist.

For the Lannan Cultural Freedom Prize I added the other recipients because it gives some insight in the relevance of the award. I did not do this for the Martha Gellhorn prize for journalism.

The point of this all is that reputation amounts to trust about the message that is written. Read the article, it is likely that you are not familiar with the Wahhabi belief, a subset of Sunni Islam that is practiced in Saudi Arabia. The article is about 200 Sunni scholars that denounce the Wahhabi belief. Several major scholars are involved. Have a read and have a think, the article is by a major journalist published in a major news paper about something that is not without consequences.
Thanks,
       GerardM

Thursday, September 08, 2016

#Wikimedia - the need for #sceptism

It is all over the news; another psychology study debunked. With two thirds of the repeated studies being debunked, there is a lot in the literature of psychology no longer valid. The source for the article I read is Mr Eric-Jan Wagenmakers professor at the university of Amsterdam.

The NWO, the Netherlands Organisation for Scientific Research, is funding 3 million Euro to repeat key research. The problem is that science is in love with what is new and quick results. Three million is at best a start.

When science cannot be relied on, collaboration with scientists and universities easily becomes controversial. The programs taught are inherently point of view and often a conflict of interest is easily established. Consider; when doctors prescribe substances that are FDA approved, it seems obvious that these substances have a positive effect on patients. Then consider that we have a Wikipedian in Residence at Cochrane, they make a reputation from debunking much of the use of such substances. We provide end user information and it seems obvious that just repeating the list of FDA approved substances without further information is not at all in our users best interest. It is even likely that we are liable for misinformation under several legislatures.

There is a need to be sceptical about sources. It is important that we not only improve the technology behind our sources, we also need an ability to mark information as debunked and have that information filter through our projects and in the information we provide. Remember, debunked is not a POV it comes with sources of its own.
Thanks,
       GerardM

Sunday, September 04, 2016

#Diversity - A Woman's hall of Fame

Wikipedia has a category of some 40 Women's hall of Fame. They are women from the past and the present that are seen as exemplary. For all the women who have an English article there is now a statement indicating that they are seen as such.

For many women who are on these lists there is no article. Obviously when the objective is to have quality articles on notable women, it is good when there are lists with articles that could be written.

There are such lists and the best thing is they is some form of automated maintenance. The Women in Red project has such lists. Many of their lists find their basis in Wikidata and it is therefore possible to add people to their lists by adding key data.

All the women who have articles are now known as such, The next thing is to add the missing articles, the red links. So far I have added items for them one by one and stated what they are known for. Obviously this is a stub. More information is needed to state what they are known for, where they lived, why they are notable. It is not only how you enrich the data it is also how you increase diversity.
Thanks,
      GerardM

#Wikidata - the conflict of interest in medical information

According to the clinical evidence handbook only 12% of the 2500 most prebscribed substances and treatments by doctors are not proven effective. There is a massive conflict of interest when unsubstantiated facts are allowed in Wikidata. Arguments like "it is NPOV" are used to defend the practice or "it is harmful for patients" when they can find out that a substance is no better than a placebo but does have negative side effects.

When an external source knows about a substance, it is fine to link to that source. This is not the same as importing the data wholesale particularly when the data is so obviously categorically problematic.

The Wikimedia Foundation has a responsibility and it is not in indicating what substances are prescribed. When we are to include information it is not on the basis that it has been approved for use but on the basis of that it is actually proven to be beneficial. An error rate of 12% on such vital information is not acceptable.
Thanks,
      GerardM

Sunday, August 28, 2016

#Wikidata - La Galería de las Mujeres de Costa Rica

#Marketing is something the #Wikimedia Foundation does not do. It does not mean that concepts like KPI are foreign to the WMF. Take this list from the English article "La Galería de las Mujeres de Costa Rica" the women listed are "women who have broken gender stereotypes and advanced human rights principals".

A lot of effort goes into fighting for a diverse Wikipedia where both women are given proper attention. If I were a marketing man, I would say that lists like this provide pointers to people who want to help. I would be happy with a list that shows all the current people with an article and I would be ecstatic when I had a list that would show all the missing articles that would auto update.

The funny thing is that technically it is not that hard to produce. It is not even that hard to include the technology into MediaWiki but it takes a marketing man to drive the point home that you have to engage people and that it shows the quality of a Wikipedia project when we know where we are lacking and where we should concentrate.
Thanks,
     GerardM