Showing posts with label consensus. Show all posts
Showing posts with label consensus. Show all posts

Thursday, July 01, 2021

What science has to say about the English Wikipedia gender gap

Why Men Don’t Believe the Data on Gender Bias in Science
 A respected Wikipedian expressed the opinion that people have it wrong when they say that English Wikipedia is biased against women. In the same week a professor stated on Twitter that her students no longer edit Wikipedia because of the toxic reception they get. As an example she mentioned a quote from an award winning scientist that was removed because "that scientist lacks relevance".

In this same week the  American sociologist Francesca Tripodi published the paper: "Ms. Categorized: Gender, notability, and inequality on Wikipedia". The paper is a scholarly read with 55 references. Most of these references are previous scholarly works, some are references to Wikipedia resources like the notability page. The references have been included in Wikidata and this is visualised in the Scholia for the paper. Please read at least the Discussion and conclusions of the paper. 

This and previous research leaves no room for evasion: English Wikipedia is biased. A personal opinion of the respected Wikipedian may differ, the consensus of the community may differ but both are biased.
Thanks,
      GerardM

Sunday, April 12, 2020

False friends and ListeriaBot - finding a way out of an impasse

ListeriaBot is a bot that maintains lists based on information in Wikidata. In this blogpost I will explain what a Listeria list is, what it is used for. I will point out its qualitative benefits and explain how Listeria can be instrumental to limit bias, stimulate collaboration and help us share in the sum of the knowledge available for us.

The heart of a Listeria list is a query. In this query it is defined what data is retrieved from Wikidata, it includes the order of presentation and shows this information in a language depending on the availability of labels.

Listeria lists are defined only once and every day a job run by the ListeriaBot updates all lists with the latest data from Wikidata. In this way available information is provided even when articles are still to be written. When there is an article to read, the label is shown in the upright position, when there is not is shows in cursive.

The biggest difference between a Wikipedia list and a Listeria list? No false friends. When you seek a specific "Rebecca Cunnigham", it is really powerful to know that your Prof Cunningham will always be known as Q77527827 and is also authoritatively known by other identifiers. From a qualitative point of view, particularly in lists, red links even blue links such disambiguation is a big thing. At this time a typical Wikipedia list has an error rate because of disambiguation issues of around 4%. I frequently blogged about this, the Listeria list I often referred to is for the George Polk award.

Maintenance is another reason to choose for Listeria lists. This was documented by Magnus, a list was maintained up to a point in time as a Listeria list and for all the wrong reasons human qualities were to prevail. Magnus compared the results after some time and the human maintained list proved to be the poorly maintained list.

Categories are lists of a kind, for many categories it is defined what they contain. Consequently Wikidata is easily updated from Wikipedias and can serve as a source for updating categories as well.

Ok, the impasse. ListeriaBot is blocked because of a false friend issue. The objective is to find a resolution that will benefit us all. The false friend issue is that images can have a same name in both Wikimedia Commons and in English Wikipedia. The existing algorithm for showing pictures is that local pictures take precedence. When ListeriaBot is to do things differently, it can. Thanks to the wikidatification at Commons, we can indicate with a Wikidata identifier what a picture "depicts". Wikidatification of images can also be introduced for pictures at English Wikipedia and it is then becomes easy to always show what Commons has unless a preference is given to show a specific image for a particular project.

I have been told that I do not assume good faith. When I see the extend people care to go to resolve this issue I am only amused. The objective of what we do is share in the sum of all knowledge and do this in a collaborative way.

English Wikipedia fails spectacularly by assuming that their perceived consensus is in the best interest of what we aim to achieve. There is no reflection on the quality brought by Listeria, there is no reflection on how its quality can substantially be improved. I fail to understand what they achieve except for feeling safe by insisting on dated practices and dated points of view.

I wish we could be one community that is known by a best of breed effort with one common goal; sharing the sum of all the knowledge that is available to us.
Thanks,
        GerardM

Friday, November 29, 2019

It is not a list when it is the result of a query

A list is a presentation of data. When a list is maintained manually, the list IS the data, when the data is the result of a query, it REPRESENTS the data.

The difference is quite important. Changing the information in a query is in the definition of the query, changing the data is a matter of re-running the query. Changing the information in a list is a lot of work and therefore there is no integrity in the data itself, it is always potluck what quality the data is.

In the Wikipedia world, Listeria is king of the queried lists. For some its use is controversial but things are changing for the better. Projects like Women in Red use Listeria a lot, their work is possible because people add notable women in Wikidata. The queries work on the basis of awards, professions, nationality enabling volunteers to write the articles they care to write. This works because once an article is written they are automagically removed from the lists.

On the English Wikipedia consensus has it that manual lists are to be preferred. However, emperically the quality of automated lists perform better {{REF}} and as data in Wikidata does not suffer from "false friends" even the support for "red links" is vastly superior.

There is no point in anecdotal evidence who is best. When the English Wikipedia has a black link for Stephen Fleming on its page for the Spearman medal first, it is an obvious start for a new item on Wikidata that is more than just a person who won the Spearman medal. It then becomes a target for lists of the special interest groups who aim to cover "their" subject matter well.

The next stage of the acceptance of lists relies on the realisation that "consensus" does not serve us well particularly when it trumps established facts. It will serve us well in politics and, in what Wikimedia projects could be.
Thanks,
      GerardM

Tuesday, September 30, 2014

#Wikipedia - #Category: #Cholmondeley Awards

A poet is proverbially poor. If there are any, it is not really an occupation but more of an aspiration to live off the high art of poetry. To celebrate great authors, great poets, organisations like the Society of Authors, recognise them with awards. Money may be involved, but the prestige, the awareness of the public is what makes the real difference for a poet.

The Cholondeley Award is to "honour distinguished poets" and, yes there is some money to be had; £8000 to be exact. There is still a category for the distinguished ladies and gentlemen who were awarded in the past. It is up for deletion because £8000 is considered chicken feed and also because awards are supposed to be in a list, not a category per WP:OC#AWARD.

It was a revelation for me that there is something like "over categorisation" in the first place. It means that categories are deleted. Wikipedia is a law unto itself and given "consensus", categories will be deleted. It is sad that all the work that went into the categorisation is deleted with those categories. It is even worse when the implied information is not first saved to Wikidata.

Categories distinguish themselves from lists because they can be found on every article that is categorised. Changes to article names are automatically reflected while Wikipedia lists are ... static. In Wikidata, lists can be defined up to a point and funnily enough this is not done for lists but it is done for categories.
The information from the category has been saved with AutoList2. Given that Wikidata often knows about more recipients of awards than any and all Wikipedias individually, the current emphasis on lists is silly. Wikidata will do a better job.
Thanks,
      GerardM