Showing posts with label Query. Show all posts
Showing posts with label Query. Show all posts

Friday, July 27, 2018

#Wikidata - I do not use query and here is why

When I edit Wikidata, I never use queries and here is why. I do not need them. For instance, I added an award to a person because it was obvious it was missing. I had no need for a query because everything that I wanted to know about the award was visible.

When you use query, you have to use a tool, define a query, run it, maybe tune it and then analyse the results. Using my beloved Reasonator, all the queries that I need are included. This is the same award and the same person but in the standard user interface of Wikidata. It is not informative, I only use it to edit.

A person wanting to teach Wikidata asked how do I structure a program? The first thing proposed was teach them to query. I agree that query is important, it has its use cases but it should not be the first introduction to Wikidata because it makes it too complicated at the start and even worse it is not necessary.
Thanks,
     GerardM

Saturday, February 13, 2016

#Wikidata - all notable #Psychiatrists - a query for big data

When you look at all the psychiatrists known to Wikidata, there are currently some 2992 psychiatrists known. When you think about psychiatry, the relevance is in the number of people who have to deal with it.

The tool used is one that does not get that much attention. It takes its time to complete but it gets whatever it is the query says. The tool does its job admirably for several years now. The one redeeming advantage is that it does what 'official query' does not offer. It can be used never mind the size of the results.

Some say that it is unrealistic to ask for this quality of service from 'official query' because "it has not been designed with this in mind". The friendliest thing to say is that this is a mistake. Official Query was supposed to replace WDQ and when it cannot by design, the design if wrong. A better argument would be that Wikidata is one of the biggest public facing resources on the Internet and people are actually using it; at this time it cannot cope. It takes money and lots of money to serve the whole world. Possibly. This approach however is an acceptable argument. It allows for seeking one or more solutions.

One existing solution is the "Toolkit", when you can have your own datastore, you can throw as much hardware at it to get results. You can give the WMF targeted money to have more hardware for you and I or implement existing software that may do the trick. We could explore if federated technology as it exists for Wikipedia could make a difference. What cannot be done is hiding behind an arbitrary choice that was insufficient from the start because official query is to replace what we already have and not take away from it.
Thanks,
      GerardM


Thursday, January 15, 2015

#Wikidata - my #bias and two articles about #diversity

As a volunteer, I spend a great amount of time making Wikidata more informative. With currently 1,896,739 edits, it is obvious that I use tools.

What I am looking for in tools is that I can use them. They do not have to be scientific, they just have to be functional. It means that I can use it at home or wherever I happen to be on any computer.

At this time there are two tools for querying Wikidata. One provides us with near real time data and the other has huge prerequisites. It is however the preferred option by people with a scientific bend.

Both approaches have been used to write about diversity. Their outcome is similar. However, I am biased towards the tool that is available to me. If I wanted to, I could run the same queries and will have have similar results. Results that will be different because of the time that has passed.

The other tool requires huge investments of me and it will only provide me with static data. Maybe the results are the same and very scientific but it will not help me improve Wikidata, it is therefore of no use to me. It reflects on data from the past. It does not compare data from the present with data I have elsewhere.

On this blog I did mention gender ratios like the two publications do. My issue with all that information is that it misses on one thing; how Wikidata is becoming more informative about diversity. As it is becoming more informative, it becomes also more useful as a tool to look at diversity in Wikipedia in the past.
Thanks,
       GerardM

Friday, November 21, 2014

#Wikimedia - first #standardisation, then #specialisation

The hardware and software used by the Wikimedia Foundation is increasingly standardised. It uses the same software and the configuration is centrally maintained. Good news; it makes for a stable platform. A stable platform allows us to share in "the sum of all available knowledge".

With this process well under way, special attention can be given to special projects. It has probably escaped your attention that the WMF now has a "Services group". They are the engineers that support the standalone software components that often run on their own machines and have very specific jobs, such as "generate a PDF from this article".

Wonderful news. When it did not escape your attention, did you notice that Stas Malyshev is getting up to speed on the Wikidata Query Service[1], figuring out what we need to do to make it suitable for widespread deployment of WikiGrok[2])?

Effectively it means that Magnus's query tool will be used by an updated version of the Games [3]. Now is that not sweet; Wikidata data being USED to leverage our community to improve Wikidata even more.
Thanks,
      GerardM
  1. https://wdq.wmflabs.org/
  2. http://www.mediawiki.org/wiki/Extension:MobileFrontend/WikiGrokhttps://wdq.wmflabs.org/
  3. https://tools.wmflabs.org/wikidata-game/

Sunday, August 31, 2014

#Wikidata - my #workflow enriching Wikidata using tools

As I have other commitments, I do not have the same amount of time to do what I used to do. The workflow I use is now quite stable and dependable so I am happy to publish it. It is fairly easy and obvious. You can do this too.

Important are objectives; mine are:
  • make Wikidata more informative by adding relevant statements
  • Provide the basis for further usage of data
My workflow is based on the people who died in 2014. This is reported in categories. ToolScript informs me about all items that do not have a date of death. Every line represents an item; typically they are human but there are also horses and other critters included. I click the Reasonator icon and, the links to articles provide me with the first lines of that article. Typically the date of birth and death are included. I copy this text when it is not English and use Google translate. From the translated text I copy the dob dod. I click on the Qnumber in the Reasonator and add these dates in Wikidata.

The ToolScript can easily point to 2013 or any other year. Obviously you can make your own script to do whatever.

Once somebody is a registered dead, I look at the article for interesting categories. They can be anything from "Alma mater university x" to "player of Whatever FC". Most interesting are the implied facts NOT reported from the dearly departed. Any category may contain hundreds of other items for whom we are not aware about said fact. The first thing to do is to document said category, this category can be on any wiki. Documenting is done by including a statement with "is a list of" "human" and have a qualifier like "alma mater" "University X". Reasonator will show at most the first 500 entries of the resulting query.

When many entries are still missing, Autolist2 is the tool to use. From the Reasonator page of the category, copy the name of the category, the P and the Q value to the appropriate spot. Do not forget to make sure that the right Wiki has been selected (en in the example). Consider the depth; depth 0 is safest. Make sure that the WDQ mode is on "AND" and press "Run". This will generate the list that is selected for processing. Check the list and copy the P and Q values to the control box. Click "Process commands" when you feel comfortable with the results. Once the process starts, you will find the changes in the Reasonator page for the item you add statements for, in the example of the illustration it is the New Zealand Order of Merit

For best results most entries are often in the "local language" like this example for people who work(ed) at the university of Innsbruck.

With a workflow like this you are more effective. The work is documented and slowly but surely Wikidata becomes truly informative.
Thanks,
     GerardM

Friday, August 29, 2014

#Wikidata - Adolf Butenandt, Nobel laureate, professor and student

For many professors we know in Wikidata that they are or have been employed by what university. Data about this has been added categories at a time. Often this has been repeated for categories about the same university from different Wikipedias.

At the same time information has been added for the universities where people studied. However, there is an increasing number of professors for whom it is not known where they studied.

Professor Butenandt is a case in point; he studied at the university of Marburg and the university of Göttingen. It is known on one Wikipedia and not on others. Given that categories are linked as well, it is fairly easy to signal missed opportunities.


Thanks to this query by Magnus, we know about 23,351 professors without an alma mater. For Mr Butenandt information has been or will be added and, obviously there is much more work left to do.
Thanks,
     GerardM


Sunday, August 17, 2014

#Wikidata - giving a #category an application

Many #Wikimedia categories have interlanguage links. Obviously the content of all these linked categories do not have the same content. Someone has to add the articles, sometimes it gets done and sometimes it doesn't. Often articles just do not exist.

When the facts that are implicit in what a category is about make it to all the items in all the categories, typically you have a superset in Wikidata. It does not stop there; items in Wikidata may be included that are not in any of those linked categories.

This is all theoretical unless ... unless you can query Wikidata and use the results. Much data has been added to Wikidata based on the content of categories and queries have been used to identify missing items this is done using AutoList2. This is one application; it is used by some of the "advanced" users of Wikidata.

What is even more interesting is showing what Wikidata things should be in a category. This is done using Reasonator. At this time for over 690 categories statements are included that define a query. This query is already complex enough that the Wikidata functionality will not be able to express the results..

These queries could be of use to "advanced" Wikipedians because it is a basis for identifying articles that have not been categorised or articles that still need to be written in their Wikipedia. For everyone else it is just interesting; this information exists and it is readily available. It is one way of learning that Wikidata knows for instance about 121,922 politicians.
Thanks,
      GerardM

Saturday, June 07, 2014

#Wikidata - Those who died in 2014


Magnus's "No date" game registers dates of birth and death. It shows that we need information for 1,258,038 of the 2,073,886 humans Wikidata knows. It is a challenge however, if Magnus proves anything it is that for a community such challenges can be met.

For the people who died in 2014, 5001 have been registered so far, it is a bit different. When people die, their articles need some attention if only to register their passing. Many of the people had their day of sporting glory a long time ago. Others like the Emir of Kano remained relevant until their end and their passing may influence many more articles.

Quality is often seen as how quickly new information finds its way in.  It is wonderful that Wikidata has the potential to flag the passing of all those known to have died in 2014 to the projects who have an article about them.
Thanks,
      GerardM

Sunday, June 01, 2014

#Wikidata - Brazilians who died in 2014 II

As an effort was made to know which "humans" are a "Brazilian, it became less difficult to know how many Brazilians died in 2014.

Currently Wikidata knows about 30,123 Brazilians, more than 10% that are "known" on the Portuguese Wikipedia. Thanks to this effort, the Wikidata number of dead Brazilians at this time is 57.

It is obvious that more notable Brazilians died in 2014. Sadly, Wikipedia does not know about them or maybe it does and Wikidata does not know that a Wikipedia does.

Relevant is that Wikidata knows about more Brazilians than any Wikipedia.
Thanks,
     GerardM

Sunday, May 25, 2014

#Wikidata - Brazilians who died in 2014

How many people from Brazil died in 2014 according to Wikidata. The answer is obvious when you know that a human is a Brazilian. So when he dies, ie a date of death is available, it is just a matter of applying the right query and you get an answer.

Surely there are more than 41 notable Brazilians who died in 2014. It means that we need to know for more humans that they are Brazilian. The category Naturais do Brasil knows about some 26,520 Brazilians, 1,128 are not known to Wikidata and 14,393 are known to be human. Of these humans 7,481 are known to be Brazilians and consequently we can safely add a statement of a Brazilian nationality to 6,912 humans.

This process is under way and the number of deaths is on the rise. Not as much as you would expect because nationality is often added when someone is registered as dead.

What is left are 1,128 articles that need an item and 12,127 items that need to become both human and Brazilian.
Thanks,
      GerardM

Monday, March 31, 2014

#Wikidata - Expressing its quality

Quality is relative. Take for instance the category about "Thai painters" or the Dutch or Thai categories. The first two know only about one painter who is Thai. The Thai Wikipedia knows about 24 painters.

At this time Wikidata knows about 20 Thai painters. It takes a little effort to add any missing painters known to a Wikipedia.

When all the Thai painters known to Wikipedia are known to Wikidata, it means that its quality to list them is as good as any Wikipedia category. Obviously there are many more notable Thai painters. They all deserve to be known in Wikidata.

Reasonator is able to express the quality of Wikidata. It uses queries that are based on the Wikidata data. By quantifying for example the painters from a given country, it becomes obvious what Wikidata has to offer. Reasonator can do a better job than any category that can be expressed as a list or a query.
Thanks,
      GerardM

Saturday, March 15, 2014

#Wikidata - #Cambridge revisited II


When you are interested in #maps, projecting the results of a Wikidata query on a map will become increasingly exciting. It seems obvious but only those items that have a geo-coordinate will find their place on a map. When I blogged about Cambridge revisited, the message was very much that we can find all items within a certain radius and, that we can process them with AutoList and WD-Fist.

Magnus is now providing us with something new. The results of a WDQ query are projected on a map. When you run a query, it could result in every item with geo-coordinates in a municipality, a county. It could be all the castles of the United Kingdom.. Just give your imagination some room.

As always, this is the first iteration of functionality. It makes use of components that we have been using for a long time. What will be interesting for us is to learn if the map functionality is able to cope. We learn by trying things out and at that, this is the perfect environment for you.
Thanks,
       GerardM

Saturday, February 15, 2014

An update to #WDQ, the tool that queries #Wikidata

WDQ is the tool where you define and run queries on the Wikidata data. It is not on official tool, it runs on replicated data in Labs and, it works really well.


While you prepare your query, you will see the number of results and the first 500 results. They are the maximum number of results that Wikidata provides when you request for information through its API.

When the query definition is satisfactory, you could create a "permalink" for future reference. What is new is that you can move on to the "AutoList:". This is where you can page through all the results and select items for special attention.
Thanks,
       GerardM

Monday, February 10, 2014

#Wikidata - These items are in Boca Raton, Florida


When you look for "Boca Raton" in the Reasonator, you will find that many of the items found are "located in Florida, United States of America". That is just dandy; when you arrive in Florida, it is obviously it is there. When you are actually in Boca Raton, Palm Beach county you stand a much better chance of finding its East Coast Railway station or its Old City Hall.

With Reasonator you find all the items that have "Boca Raton" in its name. With the "Autolist" you can query all the items that are "in the administrative-territorial entity" of "Florida". You will find many items that are in counties, cities that may be in Florida but that is NOT where you easily find it.

When you want to know that the "Boca Raton Old City Hall" is in Florida? Reasonator will show you its location just fine. It even points out that the USA is a part of our planet.
Thanks,
       GerardM


Saturday, February 08, 2014

#Wikidata - a #query for the use of #properties

Sometimes you want to do something that is just a little bit different like knowing what properties are used for the items that match a specific query.


The query that demonstrates the tool shows the properties used on the items that have "is in the administrative-territorial entity" and "Florida" as a value.

When you know what properties go together, it allows you to consider what properties are missing. A good example can be found for the combination of "instance of" and "legal case"; there are no properties identifying external legal sources in any jurisdiction.
Thanks,
     GerardM