Saturday, October 10, 2026

AI is a tool. The question is how to use it.

The sentiment among Wikimedians is clear; there is no room for artificial intelligence.

In my opinion that is stupid. You do not blame the hammer; the problem is where the hammer lands. The point of AI is in its analytics. It is not necessary to ask for new information. It can be asked to analyse existing texts, for instance all the Wikipedia articles on the same subject. The objective would be to find inconsistencies and with a corpus as large as the "sum of all knowledge" there will be plenty.

Once we are made aware of inconsistencies, the implicit challenge we always faced becomes explicit. Important is that all the knowledge necessary is in the "sum of all knowledge" itself meaning that the AI can be trained on all content from all Wikimedia projects. This would make it a multilingual and multicultural model,

What one Wikipedia knows and another does not is an inconsistency and at first the least interesting one. More relevant is to know what one Wikipedia knows different from another Wikipedia. Things like is he alive or is she the minister of defence for Whereeverland. The difference is in references. 

When an AI points to inconsistencies and to the latest references, how can any self respecting Wikimedian object to a tool like this?

Thanks,

        GerardM

Saturday, October 03, 2026

Fallout of the Language Diversity Conference

I was not at the Language Diversity Conference in Ghana and for many reasons I wish I was. I am grateful that there is a lot of messaging about events like this, results can be measured in many ways relevant for the Wikimedia ecosystem are two growth of knowledge about languages and cultures (sum of all knowledge) and the viability and relevance of languages for our global public.

At day one of the conference, Professor Samuel Alhassan Issah gave the keynote speach. One of his arguments is that translation is not enough. Relevance is gained by writing about the language, the culture the country itself that goes with a language. Anyway there is a Wikidata item for the professor it even includes a link to his ORCID identifier and that is what enables me/us to find publications and co-authors.

Adding the first publication, "The semantics of Dagbani cut and break verbs" to Wikidata brings me a few surprises. Wikidata has Dagbanli as the preferred name for Dagbani.. never mind. Also there are two co-authors both with ORCID identifiers. One of them is Bashiru Nurideen, who also has a Google Scholar identifier, I found this because I wanted to see a picture. He now has a Wikidata item and is linked to this publication. Wikidata knows about a Bashiru Nurideen Ponaa known as a volunteer, he may be the same person if he is, someone may merge the Wikidata items. Google Scholar indicates another publication for the same author. However in the publication there is no link to his ORCID identifier and his name is now Nurideen Bashiru..

In the article there is a picture who may include the professor and Bashiru Nurideen Ponaa. It would be good when there is a foto for them at Commons..

Thanks,

     GerardM

Friday, July 24, 2026

Wikimania is on and Scholia is off

Wikimania the wonderful gathering of Wikimedians is on in Paris. I have been there before, I have presented at several Wikimanias. It is truly awesome. It is the one place where all aspects of what the Wikimedia Foundation is come together, mingle and have a conversation. There is one inescapable fact; it is mostly about Wikipedia and at that, if it does not fit the English Wikipedia community, conversations will not bring you much.

The two biggest revolutions that happened lead to two new projects and apart from their function they did not change English Wikipedia. It was too expensive to have the same picture on all projects, pictures are now stored at Commons. The English community objected to its policy of only freely licensed pictures so nothing changed as they may still serve licensed materials as an illustration. Articles on the same subject were linked through "interwiki links" the process to maintain them was too costly, a centralised database was created, the links are now stable at Wikidata. The English community was not bothered because it were nerds mostly from other projects who maintained those interwiki links.

The decisions to have both Commons and Wikidata were made by the Wikimedia Foundation.

The 25th birthday of Wikipedia is celebrated in Paris and it has its challenges and opportunities. The biggest challenges in my opinion are communities that are aging and decreasing in size and information that is outdated or incomplete. Opportunities can be found in the "sum of all the knowledge that we already have". One iterative question for an AI: "What are the inconsistencies within our Wikipedias". Another is to have Abstract Wikipedia provide information for all the misses when an article is locally missing.

Why not challenge our public to be part of the solution?

Thanks,

       GerardM

PS Great tools like Scholia should be adopted by the Foundation so that it is always available and we could open up our references in a more informed fashion.

Thursday, July 23, 2026

SignWriting 2025 - A response to a Wiktionary initiative.

On Diff, the aggregator for Wikimedia news, a case was made to support sign languages particularly for Wiktionary. Sign languages are used to communicate, there are many hundreds of sign languages, only half the people who know how to sign are deaf. There is one writing system, SignWriting that can be used for all sign languages.

There are a few misconceptions in the Wiktionary proposal; with over 300 sign languages, there is a potential for over 300 translations for each concept. Sign languages are not to be linked to countries; there are countries with multiple indigenous sign languages. All languages have their own ISO-639 language code the same, ase for instance is the code for American Sign Language. 

The characters of the SignWriting script are included in the same standard that includes all other scripts like Cyrillic or Devangari. The consequence is that as your computer supports UNICODE, your computer will be able to show texts written in SignWriting when you have a supporting font. At that Google may be your friend with its Noto Sans font and the Wikimedia Foundation may provision readers of its projects content.

A decade ago, there was a concerted effort from within the SignWriting movement to have Wikipedias in multiple sign languages. Much of what transpired can still be read on my blog. At the time the Wikimedia Foundation did not have the bandwidth to support the technical requirements in MediaWiki. It is why the deaf who express themselves with their sign language have no platform within the Wikimedia Foundation.

Thanks,

        GerardM

Saturday, July 04, 2026

Growing exposure for Wikimedia content

The downward trend in the number of readers and editors of Wikipedia is a big issue but maintenance is imho as problematic. Too many articles are out of date, need an update or are factually problematic. There is not one but there are 250+ Wikipedias intended to serve the sum of all available knowledge and they all need more readers and editors to amend and append.

When we go back to first principles, references are the basis for every article and item. We could serve all known references and add reference material for people to find. They could be youtubes, scholarly papers, articles from newspapers and all that in any language. New references are linked to subjects and this makes it easier to contribute to Wikipedia articles. 

Obviously Wikipedia articles, Commons pictures, all are to be included in the search results. There is also an opportunity for Abstract-Wikipedia; even a one liner in every language is an important result particularly when it expands as more information becomes available in Wikidata.

The objective is more eyeballs for what we have on offer. "WikiFind" will expose content from partners like Internet Archive. It may expose content from YouTube, newspapers whatever.. When new content is to be exposed, it is to be linked to a topic as known in Wikidata. 

The objective is to increase our public and I may rain on the parade of many improvements considered for Wikipedia but will they improve quality or will they improve exposure. Improvements are nice to have, exposure is a must have.

Thanks,

       GerardM

Saturday, June 27, 2026

Professor Commandeur goes to Kenia - A "Wikifind" perspective

At the celebration of 25 years nl.wikipedia.org there was a game and the objective was to meet and greet fellow Wikipedians and learn about their claim to fame. I am grateful that I went as I met so many wonderful people that I had not been in contact with for many years.

Monica approached me and as we got acquainted, I learned that she was a professor of the Wageningen University. I expected that she was known to Wikidata, I was surprised by the negative, I attributed two papers to a newly minted Wikidata identifier, took a picture and added it to Commons and hey presto there was a start of a Scholia for Monica A.M. Commandeur.

We talked about "vermiculture", about African worms (they have appendages), a trip to Kenya, about ORCiD (her identifier was not public) and about Youtube. 
  • Wikidata did not know about vermiculture it was thought to be the same as vermicompost .. and then I find an item without an English label :)
  • African worms and their feet? Google is not my friend
  • Monica will travel to Kenya, we discussed why ORCiD is important; it resulted in her opening up her ORCiD identifier :)
  • The registration on YouTube of a seminar has been made public.. now for me to find them :|
In a previous post I proposed a new Wikimedia project, WikiFind. With a little effort professor Commandeur has expandeded the presence for her work. When organisations like ORCiD, WMF, CrossRef .. name them all collaborate on making freely available information findable, we would do so much better in sharing the sum of all the knowledge available to us.
Thanks,
      GerardM

Friday, June 19, 2026

Retracted and outdated sources from a Wikimedia perspective


A recent article related to the quality of Wikipedia references indicates that when a paper is retracted, the median time for a correction is 3.68 years. There is a bot for that, the RetractionBot, updated in 2024, the problem is that people have to use it AND "expecting Wikipedia editors to continuously monitor every citation for new retractions is unrealistic"...

HOWEVER

Retracted papers are hardly our only problem. Information is often superseded in later publications. This does not mean that earlier works are retracted it means that the information our articles are based on is stale. There is no bot for that AND expecting Wikipedia editors to continuously monitoring for new information is unrealistic...

ALSO

As our existing content needs maintenance, our public is diminishing and all our communities of volunteer contributors have their own objectives we have a problem; what to do?

Why not flip the script, why not provide a search engine that includes all our references, our articles and items. We enrich it with information from Retraction Watch, Crossref, ORCiD and obviously the Internet Archive. As the new "WikiFind" community adds new information items, it links them to articles and items and builds a field with potential new references. 

The objective of the "WikiFind" search engine is to be informative and provide a structure that brings our projects together AND present all of this to a new public. The implementation should be similar to how we started Commons; at the time Erik Möller started a new Wiki and only later did it serve images to Wikipedia articles... Maybe a challenge this time.
Thanks,
      GerardM