Tuesday, May 26, 2020

@WikiCommons - Meanwhile in a school in India, Japan, Russia

These students in India have to do a project. The subject is Botswana. Their teacher wants them to find many pictures so he searched Wikimedia Commons among others for pictures of  Mokgweetsi Masisi, the president of Botswana. He marked the pictures that depicts Mr Masisi and now his pupils will find more pictures of him when they look for मोकेगसेसी मासी.

At the same time in Japan students have to do a project about Botswana. Their teacher is pleasantly surprised when he find so many pictures for モクウィツィ・マシシ...
Thanks,
       GerardM

Monday, May 25, 2020

@WikiCommons - meanwhile in a different universe

And again there was a discussion that it should not be this hard to find pictures in Commons. The big difference this time is that there is now a wealth of images that have been tagged for what they "depict". They are linked to Wikidata items and they have a wealth of labels in many, many languages. In essence it has always been an objective of Wikidata to share its content in any and all of the 300+ languages supported by a Wikipedia.

The ideas that floated around soon made it into a "proof of concept" and as so often it actually worked after a fashion. The first iteration was in true Wikimedia tradition English only. The proof of concept got its second language in Dutch, Hay Kranen the developer is Dutch. Now there are nine languages and we are waiting for French to be the tenth.

So what does it do. You can look for pictures in Commons, it has 61 million media files, and when you are looking for available pictures in your language, you will find it as long as Wikidata has a label in your language.  This is for instance a result in Japanese and this is the result in German.

What can you do to make it better? Add labels in your language for the things you want to find and find media files that depicts what you are looking for. When nobody translated the software in your language, you can even do that.

Why is this so relevant? Have you ever wondered how many pictures you find in one of the smaller languages using Google or Bing? Let me tell you, it is disappointing to be polite. Commons is the repository of the mediafiles that illustrate all the Wikipedias so yes, it covers "almost anything".

The Wikimedia Foundation has this big strategy for its movement to be inclusive. This is a wonderful opportunity to show how agile it is, that it understands and supports a need that has been expressed for many many years. The beauty is the the way forward has been expressed in something that already works.

ABSOLUTELY, there will be challenges in integrating this functionality where it fulfills a need.

Luckily it is not necessary for it all to be done in one go. The first step can be as little as to take the "proof of concept" an rewrite it in the preferred language of the WMF, internationalise and localise it and keep it stand alone for now. The people who know about it will use it and they will be the first to point out what more they want to be done. A priority will be to retain its KISSable nature.

The objective is to open up Commons. Open it up in any and all languages. For me it is obvious. I will gladly give it my attention in the expectation that both Wikidata and Commons actually find a public, have a purpose that is more than what we do for ourselves.
Thanks,
      GerardM

Sunday, May 03, 2020

These scientists saw the coronavirus coming. Now they're trying to stop the next pandemic before it starts.

When you read an article with the same title as this blog post, it is one among many clamoring for attention. There is so much that can be qualified as not worth your time. In this blogpost I describe my way of adding value for articles that I think are worthwhile.

What I do is look for people in the article. In this article it is a Jonathan Epstein. The first thing is to look for Jonathan in Wikidata. Disambiguation is the name of the game and, finding candidates who might be Jonathan is the first step. Jonathan proved to be Jonathan H Epstein, there was also a Jonathan H. Epstein. Because of sharing characteristics they could be merged. Vital in this are authority identifiers and links to papers that make it reasonable to assume that they are the same person. It is helpful when Jonathan is part of the disambiguation list when people look for "Jonathan Epstein" so it is added as an alias.

The next step is to enrich the data about Jonathan P.. Authorities may identify where he works and from the website of Columbia university additional information is digested into Wikidata statements, information like the alma maters. In Wikidata many authors are only known as "author name strings", meaning they are only known as text. With available tooling, papers are linked to Q88406948, the identifier for our Jonathan.

After these steps, there is a reasonable impression of the relevance of Jonathan as a scholar and this supports the likelihood that the article that cites him can be trusted. Do this for others presented as authorities in an article and by repeating the process you provide a way for Wikidata to become a source that helps identify fake news.
Thanks,
      GerardM

Sunday, April 19, 2020

@Wikimedia interconnection, what it looks like for me

On twitter a reference was made to an article in the Sunday Times. The article is about the response of the UK government to the COVID-19 pandemic. It mentions many people and mentions their roles.

It is up to you to have your own opinion, but most if not all people are known in Wikidata, some have a Wikipedia article and all of them are in the spotlight. So when you get an edited sound bite, when you want to know if someone is "for real", it helps when you can turn to Wikimedia and find what there is to know.

This sound bite about "herd immunity" is too short to be properly understood. The argument made is that herd immunity is all that we have now that the genie is out of the bottle and, who can argue with that? Read the article as well.. After some tinkering, the Scholia for Prof Edmunds shows some 235 papers, many co-authors and still, even more co-authors are missing. The subjects he covered are extensive.. check out that Scholia. Prof Edmunds takes/tool part in UK government deliberations; it is mentioned in that Sunday Times article. He is asked to explain epidemiology to the public.

Wikimedia interconnection for me is to enrich our existing knowledge in cases like this. Tweeting about it, blogging about it may lead to even more and better information like a Wikipedia article. What we as Wikimedians do does not happen in a vacuum, connecting to what happens and who the players are help us and our readers understand who they are in  these early days of the COVID-19 pandemic.
Thanks,
      GerardM

Monday, April 13, 2020

The CDC and its National Center for Immunization and Respiratory Diseases

Because of the COVID-19 pandemic, there is so much attention to every aspect of it; the epidemiology, virology, vaccination, co-morbidity. Mix it with a heady mix of economics, profiteering and graft and what are you to think of it all. What is fact and what is not.

When I read that there is an "Outbreak Management Team" in the Netherlands, an advisory body to the Dutch government, I had a look. I added all the known scientists to Wikidata, looked for "authority identifiers" and attributed some of the papers that are likely theirs to them. It generated a really nice Scholia for them and the team as well.

At first I wanted to do similar European organisations but it takes quite some effort to find them. So I took the easy route and went for the CDC. Its organisational chart contains a wealth or smaller orgs among the the NCIRD and it has its own organisational chart. I did the same routine, adding the obvious scientists to Wikidata, looked for the authority identifiers for them, attributed papers.

The best bit? While adding people one at a time, you see how the Scholia evolves. Authors are reordered based on their number of papers, you find the ones that are co-authors and colleagues. The latest papers are shown first.. It is nice. However, this is management only, I cannot wait and see it evolve as staff finds its place in the Scholia as well.
Thanks,
     GerardM

Sunday, April 12, 2020

False friends and ListeriaBot - finding a way out of an impasse

ListeriaBot is a bot that maintains lists based on information in Wikidata. In this blogpost I will explain what a Listeria list is, what it is used for. I will point out its qualitative benefits and explain how Listeria can be instrumental to limit bias, stimulate collaboration and help us share in the sum of the knowledge available for us.

The heart of a Listeria list is a query. In this query it is defined what data is retrieved from Wikidata, it includes the order of presentation and shows this information in a language depending on the availability of labels.

Listeria lists are defined only once and every day a job run by the ListeriaBot updates all lists with the latest data from Wikidata. In this way available information is provided even when articles are still to be written. When there is an article to read, the label is shown in the upright position, when there is not is shows in cursive.

The biggest difference between a Wikipedia list and a Listeria list? No false friends. When you seek a specific "Rebecca Cunnigham", it is really powerful to know that your Prof Cunningham will always be known as Q77527827 and is also authoritatively known by other identifiers. From a qualitative point of view, particularly in lists, red links even blue links such disambiguation is a big thing. At this time a typical Wikipedia list has an error rate because of disambiguation issues of around 4%. I frequently blogged about this, the Listeria list I often referred to is for the George Polk award.

Maintenance is another reason to choose for Listeria lists. This was documented by Magnus, a list was maintained up to a point in time as a Listeria list and for all the wrong reasons human qualities were to prevail. Magnus compared the results after some time and the human maintained list proved to be the poorly maintained list.

Categories are lists of a kind, for many categories it is defined what they contain. Consequently Wikidata is easily updated from Wikipedias and can serve as a source for updating categories as well.

Ok, the impasse. ListeriaBot is blocked because of a false friend issue. The objective is to find a resolution that will benefit us all. The false friend issue is that images can have a same name in both Wikimedia Commons and in English Wikipedia. The existing algorithm for showing pictures is that local pictures take precedence. When ListeriaBot is to do things differently, it can. Thanks to the wikidatification at Commons, we can indicate with a Wikidata identifier what a picture "depicts". Wikidatification of images can also be introduced for pictures at English Wikipedia and it is then becomes easy to always show what Commons has unless a preference is given to show a specific image for a particular project.

I have been told that I do not assume good faith. When I see the extend people care to go to resolve this issue I am only amused. The objective of what we do is share in the sum of all knowledge and do this in a collaborative way.

English Wikipedia fails spectacularly by assuming that their perceived consensus is in the best interest of what we aim to achieve. There is no reflection on the quality brought by Listeria, there is no reflection on how its quality can substantially be improved. I fail to understand what they achieve except for feeling safe by insisting on dated practices and dated points of view.

I wish we could be one community that is known by a best of breed effort with one common goal; sharing the sum of all the knowledge that is available to us.
Thanks,
        GerardM

Friday, April 10, 2020

When crossing the street in the days of Corona, look left, right and left again

Many of us are at home, waiting to go out. We are all obsessed with the latest statistics and read what pundits have to say.  It is likely that you are cognizant of the statistics for your country, state or county.

I learned that Jonathan P. Tennant died in a traffic accident. When you care for statistics, you will wonder what are my chances of dying in a traffic accident at this time. Deduct it from your chances of dying of Corona and things look up.

Not so much for Protohedgehog, he met with an accident. It is sad, he was young, full of promise; just became a member of the Global Young Academy. If anything, it serves as a reminder for us to look left, right and left again to not become a bus factor.
Thanks,
     GerardM

Sunday, April 05, 2020

Edwin G. Abel aka Ed Abel

Professor E.G. Abel came on my radar because he is a recipient of the Daniel X. Freedman Award. He has a Wikipedia article as "Ed Abel" and the information of the award has him as "Edwin G. Abel".

I looked into the Freedman award because of a criticism on the Wikipedia article of Professor Montegia. The superior article of Prof Montegia is criticised because it is an orphan. It now has a Scholia template and that links the 105 scholarly papers known in Wikidata. Its timeline does include the Freedman award linking the Professors Abel en Montegia.

I doubt it is considered enough to remove the orphan template. I have added a redirect for the Freedman award to the issuing organisation. Maintaining a Wikipedia list is not one of my ambitions.. It could be a Listeria list like this one..
Thanks,
      GerardM

Sunday, March 15, 2020

#SwineFlue management with #wolves

A lot is being said about viruses and pandemics, they do not only exist in humans but also in animals particularly in kept animals. One knee jerk reaction is that by an outbreak of a disease animals in nature are blamed.

A good example is swine flue and African swine flue. It is a tradition to call for the culling, the extermination of wild boar and, traditionally the result is an increase in boar being killed.

A real solution may be found in an ecological solution, wolves who predate on boar prefer a sickly animal over a healthy animal that is better able to fight back. There is documentation of wolves determining the extend of outbreaks of a swine flue. Areas with wolves do better.

As an apex hunter the effects of wolves on its ecology are profound. There are all kinds of arguments why people oppose the reintroduction of animals that are essential for a functional ecology, animals like wild boar, beaver, wolf are extinct in places. We argue that we need more trees to offset climate change but this will not work when those trees are not placed in a functioning ecology. In Scotland trees will not grow because they will be eaten by overabundant elk.. Scotland has no functioning ecology it lacks predators like wolves and lynx to keep the elk in check.

When we consider pandemics, viral diseases, our ecology it is important to consider our own effects. We will do better when we enable ecological functionality and consider building with nature for more sustainable results.
Thanks,
        GerardM

Wednesday, March 04, 2020

@Wikipedia; the dread that is one identity that binds us all

On Twitter Janeen Uzzell praised a blogpost that is the Wikimedia Foundation All Hands: 2020 Sketchbook and indeed it informs about current thinking, most of it is great and still, I find it absolutely terrifying.

There are several great sketches in there. Katherine Maher gave an asperational talk, I love it for Wikimedia to be seen as infrastructural, inclusive and even that that what we do does not have to be in our projects. Important is that she mentions "support systems" because they provide the input for much of our processes.

Important is the page on security and risk. All the important concepts are mentioned among them; likelihood, relative impact and management preparedness but also "plan for and mitigate risks".

What truly makes me uneasy is when it is said that we aim to clarify who we are in the world in one brand, Wikipedia. The idea is that when we are all branded as Wikipedia, things are likely to become easier. When you check out the website brandingwikipedia.org there is no argument; Wikipedia is free knowledge. When you check out what it is to do
  • project and improve our reputation
  • support our movement/growth
  • be opt-in
In the abstract Wikipedia IS wonderful, in reality the concept of what Wikipedia is, is largely determined by the English Wikipedia. It it is fiercely independent, it is hardly inclusive and it has largely determined the maneuvering space the Wikimedia Foundation has. In order to "plan for and mitigate risks", I will mention several reasons why I am anxious because of this branding initiative.
  • In the Commons OTRS they use English Wikipedia notions to determine if pictures can stay or are to be removed. Commons provides a service to all Wikimedia projects
  • The query functionality for Commons is maintained by people from the Foundation. For more than half a year it puts a strain on the growth and usefulness of Wikidata. Tools have become glacially slow and often malfunction because an edit is not available when needed in further processing. It is not known what the position of the WMF director is in this
  • This is about marketing and we have never done much marketing for any of our projects. What we have done was reactive and has been all about the English Wikipedia. Now consider this:
    • Wikisource, we do not know what is available at what quality, it is all about editing and not about having people read the finished article, consequently we do not value Wikisource and fulfill its potential.
    • So far Commons has always been English only. With the support of the "depicts" functionality, there is room to enable and market  a multilingual search engine. In the spirit of "it is a Wiki", it serves as an open invite to add labels in any and all of our languages and open up what Commons has to offer. It is how to market free content the Wiki way.
    • In Wikidata we know many more concepts than what we know in any individual Wikipedias. We could use our data and inform as we have done for years in multilingual tools like Reasonator. This is an example in English Russian Chinese and Kannada. NB it takes additional labels to improve results and consequently this is the inclusive approach.
    • When Wikipedians were willing to reflect on their own performance, we could help them solve their false friends issues.
One sketch in the sketchbook is a presentation by Jess Wade. It says that even Academia is biased. As the Wikimedia community we do not need to be subservient to any bias and most certainly not the bias that Wikipedia has brought us.

Tuesday, March 03, 2020

"Building with Nature" .. a case for a beaver solution

The Markermeer is a lake with an ecological problem; the water is cloudy, plants and mussels do not grow. In order to alleviate that problem, the Marker Wadden was developed and in order to future proof the Houtribdijk the same "building with nature" concepts are used; the extensive water features will enable the growth of plants and the intended result is not only that the water will be clear again but also that the dyke will better withstand future storms.

With ecology part of the solution, it is relevant to appreciate ecology as part of a solution for open issues. There are two open issues: geese and willows. So far, geese are kept at bay at some areas with fences and young willows are being rooted out by volunteers.

When willows are allowed to grow, they will mature quickly and enable the next ecological succession. The wood and bark provides food and building material for beavers and this makes for an even more robust defense against storm damage. Some trees will mature anyway and this provides natural nesting places for white tailed eagles. Given that the wels catfish is endemic in the Markermeer, it will find its place among the Marker wadden and it may even predate on the over abundant geese.

So given that Natuurmonumenten, the organisation looking after the Marker Wadden is happy about beavers in its terrains, maybe it is the "building with nature" engineers who have to consider succession in their deliberations.
Thanks,
      GerardM

Thursday, February 27, 2020

Balancing arguments - Gender and the #Wikimedia projects

Some say, gender is important because there is a serious imbalance in the reporting on people in Wikipedia. There are many people who dedicate their time to bring some balance by writing Wikipedia articles. At the same time it is important to be cognizant of the fact that gender is not binary; the point it brings is that when you write an article you need a source to know what gender a person identifies with.

So far so good. At Wikidata other things are at play. It is vital to understand that Wikidata items are not so much about an individual, an item. When recipients of an award are included like for "Member of the Hassan II Academy of Sciences and Technologies". There is often nothing more than Moroccans that received an award because a source says so. Determining a gender relies on googling for images of the person and when the name is decidedly male like Omar, Hakim, Mustapha the gender is implied.

Why include a gender? Because projects like Women in Red rely on prospects to write articles about. Because tools like Scholia do express what we know about all the recipients of an award.. It tells us that there are currently two ladies known and 22 gentlemen. We know nothing of their work because the bias against Africa is staggering and because performance for inclusion at Wikidata is abysmal.

The arguments why we should not include gender is often based on what people expect; "Wikidata contains large sets of data and consider that it makes no statistical difference one way or the other". The reality however is that when you consider the use of data in for instance Scholia, the subsets are small. One more fine lady makes a statistical difference.

When people write about a person for a Wikipedia, they do get to know the person, they have multiple sources at hand. At Wikidata not so much. One purpose of adding people is to nibble away at our bias.

Requiring sources to indicate gender is what takes away the usefulness of the data and is counter productive when we are talking bias. For me it is a Wikipedia argument, an article based argument and it is counter productive to translate it to the set based approach of Wikidata.
Thanks,
       GerardM

Saturday, February 15, 2020

Wikipedia consensus? - It is who you ask but what are the facts

An article in VICE starts as follows: 'Wikipedia consensus is that an unedited machine translation, left as a Wikipedia article, is worse than nothing'. This article is problematic in so many ways, it starts with this premise because the Cebuano Wikipedia does not contain machine translation. It contains machine generated text and, to add insult to injury this same article states: 'the majority (generated articles) are surprisingly well constructed'.

An article like this can be sanity checked. Principles come first;
  • This is about a Wiki in contrast to the Nupedia approach. 
  • Wikipedia’s founding goal is to make knowledge freely available online in as many languages as possible.
  • There is a difference between opinions and facts
It is important how arguments are made. When "highly trusted users who specialize in combating vandalism" are introduced and comment that "many articles are created by bots", it does not follow that the quality is low nor that this is to be considered vandalism but the implication is made.

It is a fact that the Cebuano Wikipedia has 5,378,563 articles and also that there are some 16.5 million people who understand Cebuano. There is however no relation between these two facts. More relevant is that the wife of Sverker Johansson has Cebuano as her mother tongue and his two kids learn from their maternal cultural heritage also thanks to the work he does for the Cebuano Wikipedia. That is very much a classic Wiki approach.

In contrast the English Wikipedia has its bot policy preventing the use of bots for generating content. These notions should be local to the English Wikipedia and need not have relevance elsewhere. These highly trusted users can be expected to proselyte this point of view and thanks to this POV they take away a source of information without offering any credible alternative for the existing lack of information available to the rest of the world. At the same time the English Wikipedia is biased in the information it provides and does not provide the same quality of service for the domains selected for the Cebuano Wikipedia.

Sadly the Wikimedia Foundation itself makes no effective difference in support of the "other" languages it is said. An alternative to the LSJbot was introduced and it may be able to make a difference but as it does not provide a public facing service making it very much a paper tiger. Even worse are the Nupedia notions in the combination of two things: "Due to its heavy reliance on Wikidata entries, the quality of content produced is heavily influenced by the quality of the Wikidata available." and "It can discredit other Wikipedia entries related to automatic creation of content or even the Wikipedia quality.” These notions are problematic for several reasons.
  • No information is preferred over little information when our service to an end user is considered
  • Quality of information is framed in the light of existing Wikipedia entries. Whose Wikipedia entries are we considering? They are however irrelevant as our aim is to inform our end users; they do not cover the same subject.
  • When the quality is considered of Wikidata .. Why, it is a wiki and its quality is improving particularly as so many eyes shine their light on it.
  • We can inform, in any and all languages, and we do not even have to call it Wikipedia, we do not even have to save it in a Wikipedia when we only cache the results from the automated text generation.
  • When we cache results of automated text generation, texts can be generated again when the data is expanded or changed.
So far the critique of the VICE article, but then again does English not have its own problems?
  • Its 1,143 administrators and 137,368 active users are struggling to keep up, when you compare it with the 6 administrators and 14 active users for the Cebuano Wikipedia it is understandable that, as they grow, the English have to rely more and more on bots and artificial intelligence.
  • Magnus has demonstrated that the maintenance of lists is better served not by editors but by using the data from Wikidata
  • The Wikipedia technology has a problem with false friends. Arguably some 4% of list entries are wrong because the wrong article is linked to. When links are solidified by using Wikidata identifiers instead, this problem disappears in the same way as the problems with interwiki links disappeared.
The biggest problem "Wikipedia consensus" has is that it was formulated in the past by a tiny in-crowd making up the "accepted" big words for the rest of us and worse they can not be swayed from their POV by facts.
Thanks,
      GerardM

Sunday, February 09, 2020

Dear @krmaher @Wikipedia is not the #flagship to win our war

The virtues of Wikipedia have been expressed in millions of words, on many conferences and in many interviews by you, Jimmy and countless others. Nothing wrong with that. Wikipedia has been extremely useful, it has a dedicated following and it is going exactly nowhere new. What it is expected to bring is more of the same old old.

Wikipedia pundits use their own idiom, have their own values and easily dismiss what does not comply with their notions. Notions not based in actual facts but in opinions.

Our aim was to share in the sum of all knowledge, what we have is a domineering English Wikipedia expecting everything to be shaped in its image. The result is many malfunctioning sister projects that do not get attention "because what is good for the goose is good for the gander". It is not. I can find a picture of the Vasa, a former flagship but not in Commons (it uses the same technology as Wikipedia). There are many books in Wikisource but we do not know what is completed and we do not market these books to a public. When Wikidata was created its first achievement was taking inter wiki links away from Wikipedia providing a functional platform and removing millions of edits from all Wikipedias. So far functionality that does improve on what Wikipedia has is dismissed while facts show how Wikipedia under performs.

The question is if Wikimedia as an organisation is beholden to Wikipedia. If its aspirations are more than only that, it has an obligation to the other projects. It is to find a public for what the finished content of Wikisource. It has to find a public for the biggest open content resource of images making it actually easy and obvious to find pictures of for instance the Vasa. Finally there is Wikidata that is crippled by its own success and hampered from what seems to me to be a lack of organisational attention.

Dear Katherine, I am happy when a technician expresses his plans to mitigate a disaster. He does this within the restrictions he is under. It is however for you, in your capacity as director of the Wikimedia Foundation to express what relevance is given to Wikidata. We have a war chest, we are challenged to take up a new role in the war for factual and balanced information. With only English Wikipedia we have already lost the rest of the world and with English Wikipedia we also have a very biased world view. Never mind, nothing new here.

My question to you, are you aware that Wikidata has no room for growth? Is that acceptable to the Foundation? How are we going to share the sum of the knowledge that is available to us when our flagship is about to sink while sailing out of the harbor?
Thanks,
      GerardM

Saturday, February 08, 2020

The performance of Wikidata - Denis Karuhize Byarugaba

Professor Denis Karuhize Byarugaba is one of the Fellows of the Uganda National Academy of Sciences. At Wikidata we know about papers that he wrote, we know this because of the author strings that point to him.

One of the Scholia tools allows for the disambiguation of these papers by linking to the Wikidata item. It is important that we do because in this way we build on the existence of African scientists on Wikidata.

That is the theory, the practice is that it is increasingly cumbersome to even try to add papers because Wikidata more often than not informs about Too many requests. When this happens occasionally it is fine but when only 10 percent of the requests is honoured, the tool is effectively dead.

Wikidata is the most promissing tool of the Wikimedia Foundation and there is as far as I know no path forward. Obviously it affects people in what they do it affects the projects that are not progressing as fast as they could or should. Even when there is a notion of improved performance it is easily missed because of the pent up demand for much more power. Power to query and power to edit the data. We are not sharing in the sum of all knowledge when it is this hard to make it available.
Thanks,
      GerardM

Saturday, February 01, 2020

Prof Salimata WADE - some thoughts

This picture of professor Wade implies that she received multiple awards, her dress is particular to members of the National Academy of Sciences and Techniques of Senegal and multiple medals show. I added her to Wikidata but the data is sparse, it is better than what is there for most members of this academy of sciences.

In Wikidata we standardise names by having surnames at the end and they have a capital at the end. The result is Salimata Wade not Salimata WADE as you may find on many African websites..

When you google for professor Wade, it is easy to realise that she is quite notable.. It is easy even when you don't get much from French. There is work of professor Wade to find in Wikidata but attributing her work takes too much effort. She does not have an ORCiD id nor a Google Scholar ID. It is only because of googled texts that you feel safe to use quickstatements for what you find. It is super slow going but it is what you do when you expose what is possible.

Adding her papers should affect a change in two places on my African Science scaffolds, Wikipedia administrators permitting, the Listeria bot seems to be blocked for whatever reason.. Then again, other pages using the same bot are not..

When you consider the ratio of males / females it is 64 / 8. When you consider the ratio of Wikipedia articles I expect a quite different ratio. I do not know how to effectively make us a ration of US or UK scientists and compare that with African scientists. One reason is that I typically do not add nationality and I know the flaws in attributing a nationality to US scientists.. Whatever approach, Africa will show to be underrepresented both in Wikipedia and Wikidata.. Without the scaffolding, the preliminary data, there is no data approach to this.. No data means no clue.

Anyway, for countries like Senegal it makes sense to add the scaffolding to the French Wikipedia..
Thanks,
        GerardM


Sunday, January 12, 2020

Science and Africa - what colloboration exists and how do we know?

As I am adding large amount of African scientists to Wikidata, I find that I have moved into a green field. A green field as far as Wikipedia and Wikidata are concerned.

To learn about how the information about African science evolves in Wikidata, I created Listeria lists that inform about universities by country, fellows/member of academies of science and members of African young science organisations.

What I produce is a scaffolding; basic information that enables. The information that I use from the Royal Society of South Africa for its fellows includes dates, other awards, employers and even dates of death. Slowly but surely more information is being added for these people and consequently you will also find for, for instance Rhodes University, more employees and additional papers (currently only 1385 papers for its 84 scholars are known).

A scholar like Tebello Nyokong, a Rhodes scholar, has 637 papers to her name. She is a world class scientist and has four Wikipedia articles to her name. All kinds of questions may be queried for her co-authors; the gender distribution, the organisations they represent, the nationality of the co-authors.

Obviously, African science is not well represented at this time. This is a reflection of how people perceive and value African science... In essence it reflects a bias of regular Wikimedia editors. The regular Wikimedia editors are in the west, they have no reason to consider African science but this is a bias. It is highly likely that it will be hard to get Wikipedia articles accepted for African scientists because of a lack of sources and probably a lack of this perceived Western relevance.

Adding one scientist at a time does not make much of a difference. When scientists are added as part of a SourceMD process, any and all scientists who have a public ORCiD profile are likely to get included in Wikidata. This is why so many African scientist are already known. When a notable scientist is then recognised as a recipient of an award, we may already know about the papers they authored.

The SourceMD process is no longer available. It coincides with a lack of resources at Wikidata so any and all resources used for science papers are now available to something else. Understandable, but the result is that I am no longer motivated to seek ORCiD identifiers and consequently, the process is increasingly broken.
Thanks,
      GerardM

Thursday, January 02, 2020

Scaffolding in Wikidata - the Christiaan Hendrik Persoon Medal

Professor Brenda D. Wingfield is another scientist who won an award Wikidata knew nothing about. This time the Christiaan Hendrik Persoon Medal. The Wikidata basics for an award are its name, the conferring organisation and a website with details. A bonus is when there is a link to a person, an organisation the award is named after..

The Persoon Medal is a South African award, its conferring org is new to Wikidata as well. In an article about Prof B.D. Wingfield, all the other recipients are named as well; there are only six in a time span of 53 years so it was not hard to add them all.

The objective is to make connections. For both the award and Prof Wingfield, connections shows best in a Scholia. One of the more frequent co-authors is a Michael J. Wingfield.. He is a co-recipient of the Persoon Medal, a co-member of the Member of the Academy of Science of South Africa and, a duplicate of Mike Wingfield and yes, he is the spouse of Prof B.D. Wingfield as well. He is at this time Q73879566 in the Scholia waiting for papers to be attributed to the earlier item.

Another frequent co-author, Bernard Slippers, is known from a different context. He is both a member of the South African Young Academy of Science and the Global Young Academy. Given some personal connections it was easy to ask if by chance Prof Wingfield is his doctoral advisor (he is the primary, they both are).

The point of scaffolding is that it provides the structures that enable finding the data, preferably in a context. Given that most data is static, the static representation that is Listeria is really powerful. When you group them like I did for African science or Young Academies, you get the satisfaction of understanding what work is done/has been done on a subject. The icing on the cake is when you enable collaboration. I am grateful for Robert Lepenies to pick up the lead and inform other young academies for what Wikidata may mean for them. I am grateful for Daniel Mietchen for improving on the queries I use; they now show the number of publications known for each scientists and a link to the tool that enables attribution to that scientists.

My role is a simple one. I add data. Data that connects, gives relevance but most importantly data that may be picked up in queries, lists by others. The scaffolds are made by others, relevance only happens when others pick it up. My point: there has to be something to pick up.
Thanks and happy new year,
          GerardM

Friday, December 27, 2019

The value of incomplete data - Fellows of the Ecological Society of America

This is about understanding data in Wikidata. The article is about understanding what you can and cannot do with incomplete data, it is not so much about the Ecological Society of America.

The most recent work started with the news of a new Wikipedia article. Prof Cottingham is a 2015 fellow of the esa, there is a category for fellows, adding her and other missing fellows to Wikidata showed that for one fellow there was no Wikipedia article. At the time there were 90 known fellows and for only two it was known when they became a member.

I expected that new fellows would be known to Wikidata not just as an "author string" but that they would be an "item". So I added 14 of the 2019 cohort and found this not to be the case. I then looked up the known fellows from the esa webpage, added their date to Wikidata because I wondered if it were particularly the older fellows that are represented in Wikipedia.

While adding the dates, I added many alternate names to aid disambiguation, I removed one item and found two false friends; fathers mistaken for their son. When I was done, I had a good impression of the data on the website and even though I do not have the full numbers, I feel to be correct in my belief that it is the old ecology/ecologists that are represented in Wikipedia.

When you scrutinize the list of fellows, you will find included "Early Career Fellows", they are "elected for advancing the science of ecology and showing promise for continuing contributions" and they take part for a limited amount of time. Programs like these are known from all over the world and from many science orgs. This time I did not spend time on them but from previous experience I can safely say that promising is putting it mildly.

Wikidata is a wiki and as such, the work that I did is of value even though it is incomplete. I did not add all the missing fellows for instance. The esa is very much an organisation for America (check the employment of its fellows) and it takes pride in global attention and solicits membership fees from all over the world. It takes a lot of additional data when you want to compare if its subject matter is biased towards America and in what way.

For many of the fellows I added, there are papers with "author strings" waiting to be linked to an author. The same can be said for the fellows that are still missing. It can be compared to other ecological organisations but how to deal with the differences takes a completely different understanding. It takes more data to make this possible but the data does not need to be complete, that is the beauty of averages.
Thanks,
       GerardM

Thursday, December 26, 2019

Why didn’t @Wikidata have an item on Margaret Nakakeeto, a champion for living babies?

Ed Erhard wrote famously in 2018 "Why didn’t Wikipedia have an article on Donna Strickland, winner of a Nobel Prize?" A year later we can say that it is extremely likely that a Donna Strickland, a Margaret Nakakeeto are known in Wikidata if only because they are a co-author of a paper (technically: an "author string").

When Ed wrote his article, it was to highlight the gender gap we have in Wikipedia. Arguably relevant and important and it needs the attention it gets. However, it does not follow that it is the only "gap" that needs addressing, it even does not follow that the gender gap is the gap with the biggest impact.

When you consider Africa and particularly science in Africa, the subjects that matter in Africa most are reflected in for instance the Scholia for the members of the South African Academy of Science. As far as I now know, its gender ratio is 27% and this is a list with a mix of Wikipedia articles and Wikidata items. It shows the attention African science gets in Wikipedia nicely.

In Africa there is a huge amount of attention for maternal and neonatal care (eg Uganda) and as programs impact the health and survival of women, it follows that more women will become notable, notable for Wikipedia.

By giving attention to female African scientists, the subjects they are known for gain relevance. Their Scholias are developed, including links to co-authors and papers. It will improve the likelihood that when African science awards are announced, we will at least know the recipients in Wikidata.
Thanks,
       GerardM

Saturday, December 21, 2019

#Science and America first

Several US American science organisations are quite adamant that for them, it is America first. Stupidity has its place and these days the United States has a lot of it particularly as those same science organisations expect people from the rest of the world to accept "pre-eminence" of the USA.

There may be good reasons to be a member of these organisations but from my perspective, it is one thing to be with stupid, it is another to have these organisations argue their case on "your" behalf. So when you are a scientist, chances are that we already know you at Wikidata. We may even know about your science, your co-authors, your memberships.

Take for instance Prof Lise Korsten, she is probably South African, this is her Scholia. She has many co-authors and for some we do not know their gender and for most we do not know their nationality. We do not know if she is a member of any science organisation and we do not know that for her co-authors either. So you may add your professional memberships at Wikidata, your nationality and when you do know the nationality of your co-authors, you may add that as well.

In this way we make obvious to US American stupid that science is global.
Thanks,
       GerardM

Thursday, December 12, 2019

Disseminate science says @EstherNgumbi, @Wikimedia projects have the power to do just that

In this day and age science is of the utmost importance. When I am pointed to a conference where an African scientist gives the plenary lecture; the message is on display in the picture. I take an interest.

When you want to disseminate research, when you want the science to be known by society, you have to pick your platform. You can do worse than choosing for the Wikimedia projects.

Professor Esther Ngumbi is employed at the University of Illinois at Urbana–Champaign. Her ORCiD profile has only one paper but at Wikidata we knew of others. As she is now known at Wikidata with her papers, she has a Scholia. At first there was only one co-author, a bit sparse, so others were added. They were linked to the papers they have on Wikidata. The same was done for some authors who cited professor Ngumbi..

When you, your science is known in Wikidata, you are more likely to get a Wikipedia article and yes, working for an American university helps. An ORCiD profile that is open will be even more potent when you trust organisations like your university, CrossRef to update your ORCiD when it knows about your papers, your new papers.

In this day and age where our ecology is no longer stable, it is vital to know and respect the science. While we aim for the best we have to be prepared for the worst; we have to see it coming. It is why our Wikimedia projects should inform about all the science and not just what a Wikipedia article has as a reference.
Thanks,
       GerardM

Wednesday, December 11, 2019

Jack needs help, so do we and, so do our audiences

Jack penciled his aspirations for Twitter in a tweet. In it he states: "... Second, the value of social media is shifting away from content hosting and removal, and towards recommendation algorithms directing one’s attention. Unfortunately, these algorithms are typically proprietary, and one can’t choose or build alternatives. Yet."

It is good news that Jack seeks a way out, he intends to hire a "small independent team of up to five open source architects, engineers, and designers" and "Twitter is to become a client of this standard"..

In the Wikimedia projects we have similar challenges and opportunities. We cannot expect for all kinds of reasons that scientists who are very much in the news (aka relevant) there to be a Wikipedia article Dr Tewoldeberhan is a recent example but there is no reason why we cannot have her, her work and the work of any other scientist in Wikidata. With tools like Scholia we already have a significant impact by making more known that just what may be found in a Wikipedia. Jack, we do know many scientists by their Twitter handle, they already make the case for their science on Twitter. This makes it easy for you to link to and expand on Scholia. What we give our readers is more to read so that they can find conformation for what they read.

Jack, Wikidata is not proprietary, Scholia is not proprietary and the Wikimedia motto is "to share in the sum of all knowledge". Together we can shift focus from what we have read before in the Wikipedias to what there is to read on the Internet. Put stuff in context and bring the scientists who care to inform about their science in the limelight.

What we do not have is the pretense that we cover everything well. we do aim to cover everything notable well. What we provide is static, Twitter is much more dynamic and together we will change the landscape. Great technology combined with both the Twitter and Wikimedia communities has the potential of being awesome.
Thanks,
      GerardM

Thursday, December 05, 2019

What is it about Jess Wade?

It is not only that Jess writes Wikipedia articles. Others do as well. It is not only that she engages girls with science; it is why she enthuses about female (STEM) scientists. Others do as well. It is not that only that her tweets engage us with for instance the #PhotoHour, that she wants us to read the (fabulous) books by Angela Saini, she also organises for schools to have Inferior in their library for girls to read and become a scientist as well.. What makes her special is that she engages people to be part of what she communicates so well.

Take me for instance, Jess is on Twitter and I read her daily new article. For the person she writes about I enrich the information on Wikidata and ensure that the "authority control" is set in the Wikipedia article. What I add is award information, authorities, employment and education info. I often add awards and depending on how interesting an award is to me I add other recipients as well.

It is not only me, there are many more people inspired by Jess who get involved, they read the books she champions, donate so that more girls read Inferior, follow her on Twitter, write articles and also get involved, are involved. It all happens because of the enthusiasm that Jess brings to us all. This enthusiasm, the involvement is what I so cherish. When the inevitable naysayers come along it dampens the positivity, the sense that we are making a difference.

When you want to know how important the women she writes about can be, consider Joy Lawn she tweets really effectively as well... It shows how women scientists really effectively communicate the relevance of science. It is vitally important for us to know about the science, the subjects they champion. At that it may be our Jess but actually, it is Dr Jess Wade, she is a scientists first, she promotes science and Wikipedia is a vehicle to get the message out.
Thanks,
        GerardM

Friday, November 29, 2019

It is not a list when it is the result of a query

A list is a presentation of data. When a list is maintained manually, the list IS the data, when the data is the result of a query, it REPRESENTS the data.

The difference is quite important. Changing the information in a query is in the definition of the query, changing the data is a matter of re-running the query. Changing the information in a list is a lot of work and therefore there is no integrity in the data itself, it is always potluck what quality the data is.

In the Wikipedia world, Listeria is king of the queried lists. For some its use is controversial but things are changing for the better. Projects like Women in Red use Listeria a lot, their work is possible because people add notable women in Wikidata. The queries work on the basis of awards, professions, nationality enabling volunteers to write the articles they care to write. This works because once an article is written they are automagically removed from the lists.

On the English Wikipedia consensus has it that manual lists are to be preferred. However, emperically the quality of automated lists perform better {{REF}} and as data in Wikidata does not suffer from "false friends" even the support for "red links" is vastly superior.

There is no point in anecdotal evidence who is best. When the English Wikipedia has a black link for Stephen Fleming on its page for the Spearman medal first, it is an obvious start for a new item on Wikidata that is more than just a person who won the Spearman medal. It then becomes a target for lists of the special interest groups who aim to cover "their" subject matter well.

The next stage of the acceptance of lists relies on the realisation that "consensus" does not serve us well particularly when it trumps established facts. It will serve us well in politics and, in what Wikimedia projects could be.
Thanks,
      GerardM

Wednesday, November 27, 2019

Please let us support #Science at @Wikidata

When the BBC informs us about reforestation in Ethiopia.. It is Dr Tewolde-Berahan who informs BBC's Justin Rowlatt about the work that is done in preparation of planting trees.

It is a humorous piece of information that gets the message across; you can plant where trees were absent for generations and make the (local) climate change.

Consider; you now want to seriously know more about reforestation in Ethiopia. Where do you go to? Wikipedia, in all its magnificence, is rooted in its articles and thereby dated. Through its references however, there are links to its authors, to many more authors and their publications. Every article has in this way its concept cloud and it could be translated in a Scholia for an article.

The current Scholias are itself already a rabbit hole that leads in many directions and a Scholia for an article would be something different again. The article links to subjects, has its papers and by inference authors, they may link to newer papers, more papers, contradicting papers. They may lead to scientists who research similar notions for another locality.. Why not reforest Spain in France? When reforestation is possible in Ethiopia, what would be different to make this unfeasible in Europe?

And all this becomes possible when you consider Wikipedia as the jumping off point in any and all directions, not just within Wikipedia..
Thanks,
     GerardM

NB I know there are two fellows of the Ethiopian Academy of Science related to this subject. Who are they and how are they connected to Dr Tewolde-Berahan?

Thursday, November 14, 2019

@wikidata - I don't scale, help me scale

At Wikidata there is always more to do and as a volunteer you make the biggest impact when you concentrate on specific subjects. I do not scale enough to do everything I would like to do.

There are a few area's where I aim to make a difference; of particular concern is where we do not represent a body of knowledge/information in Wikidata. At this time the favour scientists particularly women, young scientists and scientists from Africa.

To make my work scale, I twitter and blog. I latch on to the great work done by Dr Jess Wade. She writes articles on well deserving scientists and I aim to add value for those scientists on Wikidata. Typically I add professions, alma maters, employers and awards. In addition I add "authorities" like ORCiD, Google Scholar and VIAF. This is important because it enables the linking of scholarly papers already in Wikidata or known at ORCiD. I can more or less keep up with Jess and, I happily add information for any and all scientists I come across on Twitter.

While doing this I learned of the Global Young Academy and as a side project started adding scientists who are member of the GYA or one of affiliated organisations to Wikidata. I am so pleased  I got into contact with Robert Lepenies. Robert is happy with the opportunity that a Scholia provides for an organisation like the GYA, for him and for all the young scientists involved. We collaborated on completing the lists on many wikipedias, Robert added many scientists to Wikidata and is now battling to keep the pictures of these young scientists on Commons...

What is crucially important for me is that Robert advocates an open ORCiD profile to scientists worldwide so that they may have their Scholia. Both Robert and I do not scale and what would help us most is an easy and obvious way that enables any scientists to start a process that will include all his papers from ORCiD, will update the known co-authors and instruct in what they can do to enrich their Wikidata / ORCiD / Scholia profile even more.

I am now working on African scientists and yes, I would appreciate some help.
Thanks,
     GerardM

PS my wife would like this scale to be enough for me

Tuesday, November 12, 2019

Instant gratification at @Wikidata

As I write this, it is 11:46am at 09:26am I added papers to prof Hafida Merzouk. The edits are picked up by Reasonator but not by Scholia. In a similar way, edits done are not picked up by Listeria.

Instant gratification is now a thing of the past, the work done at Wikidata may eventually be picked up in a Scholia or Listeria but it is not funny. Can I tweet about the things I find or have done when Wikidata no longer reflects the relevant changes?

This may sound like trivial but it does mean that when I look back at my work that  there is no longer a timely way to do so.

Instant gratification motivates and it is a factor in maintaining quality. We are losing it.
Thanks,
      GerardM

Saturday, November 09, 2019

Put (modern) #science of #Africa on the map

A young African scholar commented that the info on websites of African scholarly organisations was all about its past. There is a point to recognizing those who did good and consequently making obvious that the science of today is rooted in the past.

African scientists as well as any other scientist have a place in Wikidata with their affiliations, papers, co-authors and also with their scholarly advisors. My proposal is for all scholars to check if they are on Wikidata, check if their doctoral thesis is on Wikidata. Then add their doctoral advisor to their item and reciprocate themselves as a doctoral student.

Do not forget to include where you studied and for what university you work(ed). Check if your ORCiD profile includes trusted organisations like CrossRef that will update your profile when appropriate. When many of you do this at Wikidata you will be surprised what the impact will be.
Thanks,
      GerardM

Friday, November 08, 2019

Bias in @Wikidata and a SMART approach

When at the WikidataCon quality was presented, it was rated from 1 to 5. This approach has its own bias because it does not consider what may not be there. What is not there can be made visible using assumptions like: "a university has more than one employee" (employee includes professors) and, every country has at least one university..

The bias in Wikidata starts with the way it is mostly used and consequently how it is taught. People are shown what Wikidata looks like, immediately followed up with training in the use of query and the use of tools. At every level it takes considerable skills to make a use of Wikidata. The first hurdle to overcome is to understand the data in a single item. When your language is not English you are toast. This is Cape Town in Newari and this is a useful presentation using Reasonator. With Reasonator the information is easy to digest and adding missing labels is just one click away.

The second hurdle is knowing what bias it is you want to remedy. For a known bias like the gender gap, the Women in Red have lists of missing Wikipedia articles. A Wikidata gap is expressed by the absense of data. Listeria lists are great at that.. These are all the universities of Africa.. If you do not get the extend of what we miss, you have some thinking to do. When you apply this principle to the science of Africa, you find a lot of lists and the biggest issue remains; missing lists.

When you tackle a missing subject like I did for the "Affiliates of the African Academy of Sciences", you will find a source as a reference for the group and a reference on every affiliate. To ensure that the data is relevant and actionable, I added all of them, linked them to ORCiD and/or Google Scholar enabling SourceMD to link them to their papers. I added nationality because this may trigger inclusion on the Women in Red lists and when it was obvious, I added employers so that they may be included as a scholar on African University lists..

When we as a movement want to fight bias, we have to consider the use of lists and particularly Listeria list to show the developments of a subject. With lists available on many Wikipedias, it becomes possible to gain traction on what we miss. This approach is distinctly different as it acknowledges the need for more support for item based editing and it makes the point that missing data is a quality issue that needs to be addressed as a fundamental issue.
Thanks,
      GerardM

Thursday, November 07, 2019

@Wikipedia talks about @Wikidata

"WD is unreliable. WP:V and WP:RS are completely ignored (from any editors). International NPOV is a problem too." It is so SMART, that the best I can do is ignore it. Then again it is an open invitation to talk about Wikipedia..  There is no Wikipedia there are over 300 Wikipedia language editions.. so even the acronyms are lost on me as there is no one Wikipedia to rule them all.. 

So forget about acronyms and lets talk Wikidata and by inference raise issues particularly for the English Wikipedia where appropriate. First, Wikidata includes more items than there are subjects raised in any and all Wikipedias. Its quality can be considered in many ways and verifiability is largely ensured because of the association with other "authorities" about a subject. Thanks to the increased use of open data, it is possible to verify that specific statements are shared, increasing the likelihood that they are correct. For some information like for scientists who are a member of the AAS Affiliates Programme, we have/may have references to the authoritative source. Such references may be on a project or on an item level, it makes verifiability easy and obvious. 

Wikidata has an issue with all kinds of gaps in its coverage. For many African countries no universities are known, there are hardly any scholars associated with them. Thanks to Listeria functionality we can monitor if and when data is added. Many a Wikipedia do not have such tools because of the aversion of Wikidata by some. At the same time projects like Women in Red rely on Listeria lists and by inference Wikidata to know what to work on.

In tools like Reasonator and Listeria lists are generated and, when you compare them with Wikipedia lists, the quality is measurably better. I published frequently in the past about the Polk award.. In its lists Wikipedia has a likely error rate of six percent. When they fudge the record by not linking at all, the quality of a Wikidata lists is even better because it is much better at linking items than Wikipedia is at linking red links.  There is a solution, it just requires a willingness by Wikipedians to cooperate. 

I understand what is meant by "international NPOV" and it is where Wikidata is by definition better than an individual Wikipedia. By definition because Wikidata represents data from ALL Wikipedias. Thanks to the people of DBpedia, there is a potential to highlight where Wikipedias differ and it is more likely that the fruit of their labour will enrich Wikidata than Wikipedias.

So a Wikidatan walks into a bar..
Thanks,
       GerardM