Saturday, January 10, 2026
Hon. Erica Shafudah, Namibia's Minister of Finance and Wikimedia's sum of all know knowledge
Saturday, December 20, 2025
Maintaining information on African politicians
On many projects there are links for African politicians and when the Listeria bot is active. It will update the links whenever there is an update and it will perform this update on any Wikipedia that shares these links. For someone like Mr Bola Ahmed Tinubu it is likely that there will be an article and it will be show in the list. For someone like Mr Osagie Ehanire this is less likely and it will show a link to Wikidata.
In a personal project there are many national politicians for African countries.. The idea is that when something changes, it is reflected in the Listeria list on all the participating Wikipedias.. Recently I have done some work on Nigerian politicians, later office holders and some new Listeria lists. The new lists have to be added on other Wikipedias for them to share the latest data.
New national elections will be held in Nigeria in November .. There will be many new people who will be a candidate and compete with incumbent politicians. They all belong to parties, they studied, some may already be known to Wikidata. They may make claims about their education, their background and yes, all these claims can be verified and find a place at Wikidata. For both claims that can and cannot be substantiated there is room if only to support a public that is to make a choice.
Thanks,
GerardM
Monday, December 15, 2025
My three anwers for the questions of Bernadette Meehan
- we remain independent and thanks to our contributor communities our trust model remains in tact
- our relevance for AI training is likely to decrease because of our current inability to harness the knowledge we have and ensure validity
- when we improve the validity and consistency of our data, we will be better able to retain our English language public the main difference will be in the growth of the public for other languages
Sunday, November 16, 2025
Today's laurels are tomorrow's compost
The question becomes, how will we remain relevant and up to date. Relevancy is in multiple parts, how do we remain a challenge for our editor community, how do we remain the "go to" place for our public and how do we remain a source for the bots feeding the AI.
My suggestion is predictable. Leverage the sum of all the knowledge we have in all our projects and maximise cooperation with any and all compatible organisations.
We can share all the awards and recipients of awards known on our projects. Our academic references should all be known to Wikidata and we could and should update these in collaboration with ORCiD and CrossRef. We would have up to date portfolio for the scientists we have Wikipedia articles of. We would know for scientific articles their citations and what cited these articles. Our editors would be enabled to improve the quality of our work.
Yes, the AI engines would be better informed but hey, our intention is to share the sum of our knowledge. They are welcome to it.
Thanks,
GerardM
Wednesday, November 05, 2025
Missing award recipients in both Wikidata and the Wikipedias
Professor Fei-Fei Li is one of the recipients of the 2025 Queen Elizabeth Prize for Engineering. It says so on the English Wikipedia and it is confirmed on the website of the prize.
There are nine Wikipedias with an article for the award and there is Wikidata. When the 2025 awardees are known on a Wikipedia, "2025" should be available in the text of the article. Otherwise the article is likely out of date. The recipients should be known on Wikidata AND there should be an "award received" for the award with a date of 2025.
When you check Wikidata for this award using "Reasonator", you will find that Wikidata is in need of an update. It is by accident that I learned of this award. Updates are an hit or miss affair, this would be improved when a bot produces a list of all the awards that are in need of updates. When a bot produces this list for every Wikipedia for all the known awards, it enables people to do this maintenance work.
Obviously 2025 is this year and it will have the most mutations. A similar job can be run for other years but it is less likely to bring many additions, more likely these list will become reduced in size over time.
Thanks,
GerardM
Saturday, November 01, 2025
English Wikipedia awards, a Wikidata user story
So why not have a tool that produces a list for all awards on a Wikipedia where Wikidata knows about an award AND an award winner where both have a Wikidata item and the award winner is not on the award article. Easy obvious and it will improve the quality of articles about awards.
This can work two ways.. Why not have a tool that produces a list where awards known at Wikidata are not linked on the article.
Technically it is not that hard. It is just a few queries that are to be run on a regular basis. It is the user interface where it becomes tricky. How will a user know that something was fixed.. How will we run it for all the Wikipedias.. Will we be smart and recognise red links..
Another tool could be where we indicate to Wikipedias with an article for an award when a change happened for that award.. particularly new award winners for the current year.. It could be a list where editors are triggered to revisit their articles.
Thanks,
GerardM
Sunday, October 26, 2025
Automated updates for Wikimedia projects
I revisited my Wikipedia user page. On it I have several subpages that are regularly automatically updated when things change on Wikidata. One of them is about the "Prix Roger Nimier", I had not looked at it for years. I updated Wikidata from the data on the French Wikipedia and to make it interesting, I added the Listeria template to my French Wikipedia user page. It updated and the English and French article are nearly identical. The difference is in the description.
There are many personal project pages that are automatically updated from Wikidata. The point that I wanted to make: topics are not universally maintained. As I had another look after a few years, I found that many have had regular updates. The quality however is not that great. From a Wikimedia perspective, it seems that we have not one audience but many. When we allow for automatic updates, we will be able to share the sum of all our knowledge with a much bigger audience.
Thanks,
GerardM
Sunday, October 19, 2025
Providing Resources for a subject used in a Wikimedia project
An article typically starts with a title and it may be linked to an existing item in Wikidata. If so, the item, the concept is linked to a workflow. All the references for all articles are gathered. All relations known at Wikidata are presented. Based on what kind of item it is, tools are identified presenting information in the concept articles. They are categories and info boxes. References for content in the info boxes are included as well.
Another workflow is for existing articles. All references and relations expressed in the article show as green, unused references and relations show as orange. Missing categories and values in info boxes are presented and the author may click to include them in the article. Values in info boxes may show black, red or blue it will be whatever the author chooses.
The workflow is enabled once the concept or the article is linked to Wikidata. So for those Wikipedians who do not want to change, they just do not make use of this workflow and are left unbothered. There will be harvesting processes based on the recent changes on all projects; a change will trigger processes that may look for vandalism for new relations and for suggestions for new labels.
The most important beneficiary will be our audience. This workflow makes the sum of all our knowledge actionable to improve articles, populate articles and reflect what we know in all our articles. Our editors have the choice to use this tool or not. Obviously their edits will be harvested and evaluated in a more broad context; all of the Wikimedia projects. The smaller projects where more new articles are created will have an easy time adding info boxes and references. The bigger projects will find the relations that are not or not sufficiently expressed with references.
Providing subject resources will work only when it is supported on a Foundation scale. It is not that volunteers cannot build a prototype, it is the need for scalability and sustained performance that is not provided by the Toolforge.
Thanks,
GerardM
Saturday, October 18, 2025
Using AI for both Wikidata/Wikipedia quality assurance
All Wikipedia articles on the same subject are linked to only one Wikidata item. Articles linked from a Wikipedia article are consequently known to Wikidata. When Wikidata knows about a relation between these two articles, dependent on the relation they could feature in info boxes and/or categories in the article. At Wikidata we know about categories and what they should contain. Info boxes are known to Wikipedias for what they contain, relations are likely to be known both to Wikidata and Wikipedia
Issues identified in this way will substantially improve the integrity of the data in all our projects. We are expecting false friends and missing information in Wikidata and in all Wikipedias.
Using AI for identifying issues ensures that quality will be constantly part of the process. That basic facts are correct so that the information we provide to our audience will be as good as we have it.
Thanks,
GerardM
Monday, October 13, 2025
Batch processes for Wikidata .. importing from ORCiD and Crosreff - a more comprehensive trick
Saturday, September 27, 2025
Moving forward with Amir's "Internal Links in #Wikipedia" presentation
Functionally, every link red or blue should remain exactly as is. Technically, every blue link refers to one article and every article SHOULD have an item at Wikidata. Every link, blue or red, may be referred to from many places and SHOULD be about only one concept. For every destination there MAY be a link to an item at Wikidata. At this time we have no way of knowing if there is only one concept and if there is an item at Wikidata for that concept.
Many years ago Wikidata solved a similar problem. Wikidata was an instant success because it replaced the interwiki functionality. The solution proposed today is similar and only possible now that Wikidata can be "federated" with many instances of a Wikibase.
All destinations for both red and blue links will be known in a local Wikibase federated with Wikidata. Any destination may be linked to a Wikidata item but the name of the local article/destination will remain unique. Thanks to this federation, disambiguation support may be provided based on what is known both locally and globally when a new link is created. It will know about the synonymy for each subject.
This change does not need to be controversial because like with the interwiki links, people can opt out of this new functionality. When only a subset of the editor community becomes involved, the quality of all links will improve quickly. With the interwiki links fixed, Wikidata was ready to become a knowledge base. As the wiki links in the local Wikibases get in shape, the Wikidata knowledge base may be used to signal that articles should be in specific categories, or that red links could be added in summation articles like in articles about an award.
Our dependence on Wikipedia editors will remain key but tools like the Wikidata knowledge base are available to bring us the data that enables us with information that is up to date and improves the connections between all our articles. Manually checking wiki links is a Sisyphean task, with tooling it becomes manageable and worthwhile.
Thanks,
GerardM
Batch processes for Wikidata .. importing from ORCiD and Crosreff
This was done in the past by a different tool. It was a drama because Wikidata is NOT a relational database. The problem is that an item cannot be created with the certainty that it will be unique. To ensure that new items will be unique there are plenty of available tricks.
The easiest trick is to have an option in the tool to create all the missing papers known for a given author. One author at a time and, from Scholia. It makes use of results from a batch process that runs once a week. Cheap, cheerful highly effective.
Then there is a need for another batch process. For all the "author string"s that include an ORCiD identifier, existing authors are sought and these author strings are changed into "author"s removing the link to the ORCiD identifier as it is implicitly part of the author. This process can run once a week.
A second batch process, also running once a week, looks for "author string"s with ORCiD identifiers without corresponding authors. It generates a list of ORCiD identifiers with associated "author string"s and creates one new item uniquely identified by that ORCiD identifier.
Obviously new authors make it useful to run the first batch process again.
These batches could run exclusively for an author processed by Orcid-scraper making this tool and Scholia more powerful and up to date.
Thanks,
GerardM
Sunday, September 14, 2025
One line in a Wikipedia article; a prize is name after her
This award currently has no Wikipedia article, it has a Wikidata item and consequently associated information can be shown in Reasonator or in a Scholia.
I added the 2024 recipient because awards without recipients is not really informative. I came across the inaugural 2013 recipient because of another award she received. All six recipients had an Wikidata item and only one did not have publications associated with him. However, a merge of two items solved that issue.
It is wonderful to see prestigious organisations refer to Wikipedia articles. I do notice that we are still at a stage where Wikipedia information is not valued enough to mine it, curate it and finally share it with all our audience. We could share the knowledge that is available to us.
Thanks,
GerardM
Sunday, August 31, 2025
3 million is a lot of Wikidata edits
Wikicite would be a big thing when it has room for growth. It should contain all the references used in Wikpedia, it could contain all the awards known to Wikipedia and all its recipients. It would be great when all publications referencing Wikipedia sources would be known; it would be a clue to how much Wikipedia readers are missing out on.
When we know about all the publications of the CDC scientists at Wikidata, thanks to a Scholia presentation we would know what America has developed so far. It is a lot, we could celebrate it.
This is the first blogpost for me for 2025. I still dream big but what I achieve is one edit at a time.
Thanks,
GerardM
Thursday, November 14, 2024
Red pill and blue pill - Wikipedia is it a binary choice?
As far as the English Wikipedia is concerned, there is no red nor a blue link for the 2024 awardees of the Brewster medal. Its information ends in 2021. The German Wikipedia is up to date. There are no articles for Renée A. Duckworth and for Juan C. Reboreda on both Wikipedias, the German has two red links.
When you maintain information like this, there are three options. You can include an awardee in text or as a link and as luck will have it the link will turn red or blue. This is complicated because a link may have homonyms. With a red link you will only know an homonym issue once an article is created, with a blue link you may know immediately.
The Wikimedia Foundation solved a similar problem a long time ago for another type of link, the "interwiki link". The solution is Wikidata. It works because there is only one identifier for every topic and every article needs a link to a Wikidata item to have a more global relevance.
Thanks to the ongoing development of Wikidata, there is the Wikibase. We should do a similar job for the red and blue links. It will do away with the false friends problems in Wikipedia. It will improve quality for each Wikipedia and it will improve the quality of Wikidata. Any data related updates that are not strictly local will remain at Wikidata because that helps us in the sharing of the sum of all knowledge.
When a new a link is to be added in any of the 333+ Wikipedias, it starts with disambiguation.. Is the subject already known in any of the other Wikipedias? If not a new Wikidata item will be created and extend options in any future disambiguation. If it is, available information and references are available from the start and consequently a Scholia, a Reasonator or any other generated view of the information may become available dependent on the policies of a Wikipedia.
Implementing such a Wikibase is not really problematic because all the blue links still refer through the local Wikipedia article to Wikidata. The red links are the more tricky bit. They are opened up once they are linked to a Wikidata item.
With such a Wikibase in place, we can start doing the smart things. The Brewster medal, Q612041, could have a red or blue link to all the awardees. When they don't the article is to be reported for maintenance..
Cool?
GerardM
Tuesday, November 12, 2024
Fellows of the Royal Zoological Society of NZW and .. #ChatGPT
Wikidata did not know the award.
The list of fellows on the RZS website is formatted in a "last name, first name" format. There are too many fellows so converting it by hand is inconvenient. As so many people are enamoured by ChatGPT, I gave it a spin. ChatGPT does NOT process websites for me. So I copy pasted the list and asked it to change the order of the surname and the first name.
I asked it who had a Wikipedia article. It could not tell me but it gave me a list of fellows who likely have a Wikipedia article. For many of them I added the award in Wikidata and for some fellows I added a new Wikidata item. For many of them I linked publications and this results in a nice Scholia for the award.
It would be really cool when there is a Wikimedia AI that will answer questions like: "for the people in this list change the order of the name and check if these Australian award winners have a Wikipedia article or a Wikidata item". Maybe start with a tool for editors and then open it up to the general public.
Given that Wikipedia is multilingual, what would be the effect of the data for the answers being all Wikipedias AND Wikidata.. Given that Wikifunctions is language agnostic, why not have functions that are a front end to such a Wikimedia AI?
Thanks,
GerardM
Saturday, November 09, 2024
The story of African award winning scientists using Wikifunctions
You can find the winners of the Alan Pifer Research award on the English Wikipedia. One of them, the 2011 recipient is Mr Kelly Chibale. there are several ways to be informed about him. There is Scholia and Reasonator, both derive from Wikidata and then there is the Wikipedia article. All four provide information, one is unstructured and exclusively in English. The good news is that parts of it have a structure making it easy for tools to analyse and convert to data.
A person can read an article, find and add an award not in Wikidata and choose to add the awardees or use "Awarder" to do it with less effort. It is good when it is done but analytical tools could do a better job. There are many tools that produce information in a nice layout like Listeria.. Problem is that it is not maintained by Wikimedia and it is not necessarily multilingual.
And then there is Wikifunctions. It is developed and maintained by the Wikimedia Foundation. It could do all the things that Listeria does. Having a function that does only list all the honours and awards for someone like Mr Chibale would be great particularly when there is a function that brings to the light all the award winners for any award. An article about an award can be minimalist, and still include stuff that typically goes into an info box.
With functions available like this, it PAYS to engage in Wikifunctions for the specifics of a language for a function. It is impossible to include all awards in any language but with some imagination, we can expose information once the necessary functions are available.
Thanks,
GerardM
Tuesday, October 29, 2024
the virality of co-authors in urology
Happy birthday Wikidata and, many happy returns.
When you start enriching the data for a Dutch urologist, an academic who published quite a number of scientific papers, obviously there must be many co-authors. Many of them are yet to be identified, at this moment for Jakko A. Nieuwenhuijzen there are some 339 still to be added.
The main consideration is what has the biggest impact. As a colleague of Mr Nieuwenhuijzen is known at Google scholar, adding papers for him brought new publications to Mr Nieuwenhuijzen and many of his co-authors. Enriching data for these co-authors makes the graph more complex.
At some point more precision in the data for a single author is no longer worth the effort. When you then find an other urologist with many papers not yet attributed and many co-authors where Wikidata does not know the gender yet, focus shifts and many more edits make their way into Wikidata.
Many of these co-authors are of the same institute but people from elsewhere find their place in these graphs as well. Many are Dutch but as urology knows many international collaborations this is reflected in the expanding number of co-authors.
As a topic is developed in this way, it easily results in thousands of edits. As many subject are researched in this way, the enriched data is there for the world to use. This data is only of value when there is a public. Sharing in the sum of all knowledge has always been what we stand for. Sharing freely and widely generates us a a both public and a future.
Thanks,
GerardM
Sunday, October 27, 2024
The fallibility of notability
When Wikidata will be split up in a "science" part and "all the rest", scientists who have a Wikipedia article will need to be part of the "rest" as well. This is necessary as all Wikipedia articles have a link to Wikidata because of the "interwiki" mechanism.
It follows that there will be an over abundance of USA scientists and there will hardly be any scientists of Africa or South America.
Some data about scientists is likely to be considered to be part of "all the rest" awards for instance. Are these scientists who received an award to be known in two data sets? Some scientists had a career as an athlete.. an other reason for duplication. It is hard enough to maintain the interwiki links and existing duplication within Wikidata, it will become exponentially more difficult when another data set is added.
When the creation of Wikimedia Commons was considered, similar good reasons led to hesitation and prevented us to bite the bullet for quite some time. Commons started with the creation of a Wiki, a MediaWiki patch that showed a picture in a Wikipedia and it then took a long time for most of the duplicate pictures to be only in Commons. It was not technically perfect but it was done perfect in the wiki way.
I hope that we will bite the bullet this time as well. With a new unrestricted wikibase, the old batch jobs can be dusted off and make good for the years of academic data we missed. I pray that Scholia will become functional soon after.
I will still be able to do my Wikidata thing.. projects like African politicians, Muslim countries and their rulers (past and present).. Awards that can do with an update obviously including science awards.. I will not be bored but maybe I will be working .. maybe not.
Thanks,
GerardM
Saturday, October 26, 2024
Old soldiers never die, they march in the remembrance parades
As our movement matures, people who were there at the beginning, age. They get other priorities, they get sick, operated upon and as a consequence have a windfall of time to do more work at Wikidata.
I did a similar job for a dear fellow Wikimedian.. It is now my turn, my chirurg is in this picture and as I add missing co-authors this picture becomes more complex. It will also become more complex when existing co-authors are enriched with new and linked papers.
With Wikipedia there is the promise that even though the information will evolve, all the work people have put in will be there in future and enable people to read/study the subjects each editor cared for.
The data of Wikidata as it is will be split in parts. For the best of reasons but once its structure is broken, the tools that bring structure to the data will be broken as well. The same tools that enable the enrichment of the data will be broken. Much of my Wikimedia legacy will be lost because there will no longer be a public enabled to learn about scholarly works in a Wiki way.
For a few years now this sword of Damocles has hung over Wikidata. As a consequence the potential of Wikidata is not being realised. The data could be so much richer when automated processes bring free knowledge together. References in Wikipedia indicating later papers and improve its quality.
As long as I can I will do my Wikidata thing; hope is eternal.
Thanks,
GerardM














