Saturday, March 09, 2019

A #marketing approach to "what it is that people want to read in a @Wikipedia"

All the time people want to read articles in a Wikipedia, articles that are not there. For some Wikipedias that is obvious because there is so little and, based on what people read in other Wikipedias, recommendations have been made suggesting what would generate new readers.This has been the approach so far; a quite reasonable approach.

This approach does not consider cultural differences, it does not consider what is topical in a given "market". To find an answer to the question: what do people want to read, there are several strategies. One is what researchers do: they ask panels, write papers and once it is done there is a position to act upon. There are drawbacks; 
  • you can only research so many Wikipedias
  • for all the other Wikipedias there is no attention
  • the composition of the panels is problematic particularly when they are self selecting
  • there are no results while the research is being done
The objective of a marketing approach is centered around two questions: 
  • what is it that people are looking for now (and cannot find) 
  • what can be done to fulfill that demand now
The data needed for this approach; negative search results. People search for subjects all the time and there are all kinds of reasons why they do not find what they are looking for.. Spelling, disambiguation and nothing to find are all perfectly fine reasons for a no show. 

The "nothing to find" scenario is obvious; when it is sought often, we want an article. Exposing a list of missing articles is one motivator for people to write. Once they have written, we do have the data of how often an article was read. When the most popular new articles of the last month are shown, it is vindication for authors to have written popular articles. It is easy, obvious and it should be part of the data Wikimedia Foundation already collects.. In this way the data is put to use. It is also quite FAIR to make this data available. 

For the "disambiguation" issue, Wikidata may come to the rescue. It knows what is there and, it is easy enough to add items with the same name for disambiguation purposes. Combine this with automated descriptions and all that is requires is a user interface to guide people to what they are looking for. When there is "only" a Wikidata item, it follows that its results feature in the "no article" category.

The "spelling" issue is just a variation on a theme. Wikidata does allow for multiple labels. The search results may use of them as well. Common spelling errors are also a big part of the problem. With a bit of ingenuity it is not much of a problem either.

Marketing this marketing approach should not be hard. It just requires people to accept what is staring them in the face. It is easy to implement, it works for all the 280+ language and it is likely to give a boost to all the other Wikipedias but also to Wikidata.
Thanks,
        GerardM

Sunday, February 17, 2019

@WikiResearch - Nihil de nobis, sine nobis

There is this wonderful notion how Research is going to tell us what to do in light of the strategic Wikimedia 2030 plans. Wonderful. There is going to be this taxonomy of the information we are missing.

Let me be clear. We do need research and the data it is based on, it is to be available to us. There is no point in a future taxonomy of missing knowledge when we have been asking for decades : "what articles are people looking for that they cannot find". If there is to be a taxonomy what else should it be based on?

When we are to fill in the gaps of what Wikipedia covers, we can stimulate more new articles by indicating what traffic they get in the first month. Stimulate our readers to learn more by showing what Wikidata has to offer and show its links to texts in other languages. It may even result in new stubs even articles in "their" language. This technology has been available for years now.

The WikiResearch is full of arguments on the importance of citations and Wikidata as the platform for all Wikipedia sources, why then are the WikiResearch papers not in Wikidata from the start. What is it, that WikiResearchers consider that Wikidata is not about them? Just as it is about any other subject Wikidata covers? What is it that makes their work less findable (FAIR) than what is known to have been published as open content by the NIH?

The point I want to make is that no matter how well intended it is what the WikiResearch aims to achieve, they lose the interest, involvement and commitment of people like me, the people they need to get the results they aim for.

Yes do research, but we should not wait for its results, we know how to stimulate people to write new articles.
Thanks,
      GerardM

Sunday, February 10, 2019

#Wikidata - A quick and dirty "HowTo" to improve exposure of a subject in Wikidata

When you want to expose a particular subject, any subject, in Wikidata. This is the quick and dirty way to expose much of what there is to know. There are a few caveats. The first is that the aim is not to be complete, the second that it is biased towards scientists who are open about their work at ORCiD.

You start with a paper, a scientist. They have an DOI / ORCiD identifier and, they may already be in Wikidata. First there is the discovery process of the available literature and the authors involved. The SourceMD tool is key; with a SPARQL query or with a QID per line, you run a process that will update publications by adding missing authors or it will add missing publications and missing authors to known publications.

When you treat this as an iterative process, more authors and publications become known. When you run the same process for (new) co-authors, more publications and authors become known that are relevant to your subject.

To review your progress, you use Scholia. it has multiple modes that help you gain an understanding of authors, papers, subjects, publications, institutions.. You will see the details evolve. NB mind the lag Wikidata takes to update its database. It is not instant gratification.

A few observations, your aim may be to be "complete" but publications are added all the time and the same is true for scientists. People increasingly turn to ORCiD for a persistent identifier for their work. The real science is in designating a subject to a paper. Arguably the subject may be in the name of the article but as an approach it is a bit coarse. I leave that to you as your involvement makes you a subject "specialist".
Thanks,
       GerardM

Tuesday, February 05, 2019

#Wikidata - Naomi Ellemers and the relevance of #Awards

In a 2016 blogpost, I mentioned the relevance of awards. At the time Professor Ellemers received an award and it was the vehicle to make that point in the story.

Today in an article in a national newspaper, Mrs Ellemers makes a strong point that the perception of awards is really poblematic. What they do is reinforce a bias that American science is superior. It leads to a perception by European students that it is the USA "where it is all happening". A perception that Mrs Ellemers argues is incorrect.

NB Mrs Ellemers is the recipient of the 2018 Career Contribution Award of the Society for Personality and Social Psychology.

Wikidata re-inforces this bias for American science by including a rating for "science awards". This rating values awards by comparing them. This rating is done by an American organisation and the whole notion behind it is suspect because the assumptions are not necessarily / not at all beneficial for the practice of science.

How to counter such a bias? As far as I am concerned there is no value in making a distinction between awards and "science awards" and biased information like this should be removed. Just consider, when European science is considered less than American science... how would  African science be rated?
Thanks,
     GerardM

Sunday, February 03, 2019

Dr Matshidiso Moeti, an exeption to my rules

When I add scientists to Wikidata, I really want something to link to, an external source like ORCID, Google Scholar Viaf.. When I link publications it is the data at ORCID I link to, I don't do manual linking.

From the sources I have read, Dr Moeti is the kind of person who deserves a Wikipedia article. Her work and the people she works with, the cases she works not only deserve recognition it is imho vitally important that they do, that you learn about them. This is why I made exceptions to my rule.

This is her Scholia, this is her Reasonator and please, take an interest.
Thanks,
      GerardM

The case for #Wikimedia Foundation as an #ORCID member organisation

The Wikimedia Foundation is a research organisation. No two ways about it; it has its own researchers that not only perform research on the Wikimedia projects and communities, they coordinate research on Wikimedia projects and communities and it produces its own publications. As such it qualifies to become an ORCID Member organisation.

The benefits are:
  • Authenticating ORCID iDs of individuals using the ORCID API to ensure that researchers are correctly identified in your systems
  • Displaying iDs to signal to researchers that your systems support the use of ORCID
  • Connecting information about affiliations and contributions to ORCID records, creating trusted assertions and enabling researchers to easily provide validated information to systems and profiles they use
  • Collecting information from ORCID records to fill in forms, save researchers time, and support research reporting
  • Synchronizing between research information systems to improve reporting speed and accuracy and reduce data entry burden for researchers and administrators alike
At this time the quality of information about Wikimedia research is hardly satisfactory. As is the standard; announcements are made about a new paper and as can be expected the paper is not in Wikidata. The three authors are not in ORCID, as is usual for people who work in the field of computing so there is no easy way to learn about their publications.

What will this achieve; it will be the Wikimedia Foundation itself that will push information about its research to ORCID and consequently at Wikidata we can easily update the latest and greatest. It is also an important step for documentation about becoming discoverable. It is one thing to publish Open Content, when it is then hard to find, it is still not FAIR and the research does not have the hoped for impact. It also removes an issue that some researchers say they face; they cannot publish about themselves on Wikimedia projects. 

Another important plus; by indicating the importance of having scholarly papers known in ORCID we help reluctant scientists understand that yes, they have a career in open source, open systems but finding their work is very much needed to be truly open.
Thanks,
       GerardM

Sunday, January 27, 2019

@Wikidata #quality - one example: Leonardo Quisumbing

Quality happens on many levels. Judge Leonardo Quisumbing passed away and a lot of well meant effort went into his Wikidata item.  The data is inconsistent with our current practice so in the Wikidata chat people were asked to help fix the data.

Judge Quisumbing held many positions, one of them was "Secretary of Labor and Employment". This is a cabinet position and it follows that Mr Quisumbing was also a "politician". It is one thing to include this position and occupation to a person, from a quality point of view it is best to include a "start date" a "replaces" an "end date" and a "replaced by". The problem: the predecessor and successor do not exist in Wikidata.

Many a secretary of Labor do have a Wikipedia article and they are included in a category. Using the "Petscan" tool it is easy to import all those mentioned. Typically the quality of the info is good however there is always the "six percent" error rate. Indeed one person was erroneously indicated as a "secretary of labor".  The problem is that people who only care about quality on the item level are really hostile to such imported issues. They are best ignored for their ignorance/arrogance.

A next level of quality is to complete the list with all missing secretaries. This can be done warts and all from the Wikipedia article. It results in a Reasonator page that includes all the red and black links of the article. Many new items are created in the process and having automated descriptions are vital in finding as many matches as possible.

Judge Quisumbing became an "Associate Justice of the Supreme Court of the Philippines" and became the senior associate justice in 2007. Adding associate judges from a category was obvious, adding senior associated judges is a task similar to secretaries of labor. However, a senior is the first among the many and consequently it requires a judgment call on how to express this.

Given that Wikidata is a wiki, you do the best you can to the level that has your interest. There is still a need to improve the Wikidata item for judge Quisumbing but that is for someone else.

Thanks,
       GerardM

Sunday, January 20, 2019

@Wikidata - #Quality in a #Wiki environment

What quality is, quality in a data environment has been studied often enough. Lots of words are spend about it but one notion is always left out. What is data quality in a Wiki environment. How does that translate to Wikidata.

First of all; Wikidata serves many purposes. The initial purpose of Wikidata was to replace the in-article "interwiki" links. They were notoriously difficult to maintain, often wrong. A single Wikidata item replaced the links for a subject in all Wikipedias and this brought stability and a high level of confidence in the result. Over time the quality of the "interwiki' links went down; there are fewer people involved adding and curating these links and it is seen as a quality issue when new items are generated for new articles; they do not have statements and are often not linked. There have been protests against these new additions.

A second purpose is the use of Wikidata statements in Wikipedia templates. Assessing data quality becomes complicated as there are micro, mesa and macro levels of quality at play. The micro level: is sufficient data available for one template in one Wikipedia article. The mesa level: is sufficient data available for one template in Wikipedia articles on the same topic. The macro level: is the same data available for all interested Wikipedias and do we have the required labels in those languages.

Quality considerations are driven by this approach. On a micro level you want all awards for a scientist to be linked on an item. On the mesa level you want all recipients of an award to be linked to their item. On the macro level you want all awards to have labels in the language of a Wikipedia and have all local considerations been met.

Standard quality considerations in a Wiki environment are not helpful; they are judgemental. People contribute to Wikidata and all have their own purposes. A Wiki is a work in progress and when quality assessments are to be performed, the question should focus on the extend a specific function is supported. What people seek in support also changes; as long as there was no article for professor Angela Byars-Winston it was fine only to know about her for one publication. Now that Jess Wade picked her for an article, it may be relevant that she is the first and so far only person known to Wikidata who was a "champion of change" and that more papers are identified for her.

Wikidata includes many references to scientific papers and authors. However, so far it serves no purpose. Allegedly there is a process underway that imports papers used as citations in the Wikipedias but it is not clear what papers are used in what Wikipedia article. So far it is a big stamp collection, a collection with a rapidly growing quality. A collection that highlights authors who are open about their work and who share the details of their work at ORCID. In effect, this data set indicates that the relevance of a scientist improves by being open.

Wikidata invites people to add/curate the data that is of interest to them. Particularly the esoteric data, data about subjects like African geography, Islamic history need a lot of tender loving care. It is where Wikidata and the large Wikipedias are weak. For as long as Wikidata is largely defined by the large Wikipedias it will reflect the same biases and these biases will be hard to assess and curate.
Thanks,
      GerardM

Tuesday, January 01, 2019

The #decline of #Wikipedia (as we know it)

Regularly, we are told about misgivings about Wikipedia. It can not stay as it is, it is in decline; it is all doom and gloom.  NB the use of the phrase "doom and gloom" increased in the 1950s.

So Wikipedia will not remain as we know it? GOOD, it forces us to think how we can improve what we have. When things are to change, what will have a healthy impact? How will we get something that serves us better in "sharing the sum of all knowledge". How will we get more people use what we have to offer and how will we entice more people to contribute to the data collection that is included in all the Wikimedia Foundation projects.

First thing; our projects need to be less US-American. For me, a POV situation I was in, was "obviously"decided in favour of only considering the USA point of view; I let it slide but went to pastures green. The money we raise is for: "keeping the servers going". An objective a bit too limited to my taste but it raises the cash. Money is mainly raised in the USA but in order to be truly global, it is better to raise more equally in every country at least for the amount it cost to serve it. Gapminder is where you may be reminded that money is everywhere. As to the servers, why have all crucial eggs in one USA basket? Given its current politics, there is indeed a potential doom and gloom scenario possible. Having them more dispersed will bring our data closer to our audience, our editors as well. Benefiting them with better performance; that is the easy win. A more complicated solution is in the implementation of the Vrije Universiteit research of a peer to peer MediaWiki.

When our projects are to be less US-American, it is important for spending to be more global too.

When today's Wikipedia practices are no longer considered to be set in stone, we can finally implement features that enable, ensure and enhance its future. First, we should be less self centric; after all there is only one sum of all knowledge and we define only a part of it. Magnus showed how to maintain lists in an efficient way and Amir added recently a "task" to Phabricator to implement proper disambiguation of "red links". We are increasingly aware, not only of the references of all Wikipedias but also of publications by scientists that enable their work to be found. Complement this with the scientific papers we publish and we improve the public relevance of scientists by making them findable, by pointing to their science.

With a changed approach at Wikipedia, we may be bold and change the outlook on what Wikipedia is there for as well. Why not make Wikipedia the gateway to information held elsewhere? Why not show a Scholia page for every scientist we know, why not offer the books at OpenLibrary or inform on the availability of books at the local library?  Why not partner with other organisation we have a shared objective in. But most importantly let us be aware that an African professor teaches in Africa and that we allow for and enable the context of our partners and volunteers.

For me there is no reason for doom and gloom as there are so many opportunities to become even more effective. With a whole new year in front of us; let us do well.
Thanks,
        GerardM

Wednesday, December 26, 2018

Professor @steve_hanke and reading what is #FAIR

Professor Hanke is on Twitter. He has his five Wikipedia articles and his info on Wikidata is well developed. With a scholar of his eminence, you would expect a lot of known publications as well. However, never mind the 153 English Wikipedia references, never mind the links to 13 external authorities, finding his work is not easy nor obvious.

The problem with Wikipedia references, it is a hodgepodge of links about him and links to his works. His VIAF registration may bring you some of his works but it will not tell you where his books are cited. Mr Hanke does not have an ORCID identifier and consequently it is not easy to include his data on Wikidata.

This is not about Mr Hanke; in certain fields of science people do not have an ORCID identifier or are not open about their publications. When you are interested in a specific subject or a specific scientist, it helps when the information is FAIR.

So what is missing; there is this database with all Wikipedia references, it needs to be included in Wikidata as soon as possible. It may require a fair deal of social manoeuvring to include all Wikipedia references to Wikidata. But the benefits; the benefits will be huge. Given that Wikipedia references are backed up by the Internet Archive, this will extend for these links in Wikidata as well. It makes them FINDABLE and ACCESSIBLE. At Wikidata, this data becomes INTEROPERABLE and REUSABLE (FAIR).

So my 2019 wish for the Wikimedia Foundation is to become FAIR in what it says and what it does.
Thanks,
      GerardM

Tuesday, December 25, 2018

Dear Katherine: Socialization Tactics in Wikipedia and Their Effects

In the contract of Wikimedia employees it says that they are not allowed to blow their own horn in any of the Wikimedia projects. It is according to a very senior Wikimedia official why they cannot add/contribute to information to scientific papers like Socialization Tactics in Wikipedia and Their Effects in Wikidata.

Dear Katherine, you will agree with me that this is a perverse effect of a well intentioned item in personnel contracts. So let me tell you more about the effects and how we can overcome this issue.

As you know, there is a thriving research community and its recorded presentations showcase the  research on Wikimedia projects. These presentations are recorded and may be found on YouTube. Typically these showcases are based on scientific papers. They should be recorded in Wikidata with all the details like it is done for any and all subjects. When a paper is properly covered, we know all its authors, the papers it cites and in time the papers who in turn cite the paper. When an author is well covered, we know every paper published, co-authors, subjects, subjects, citing authors. We know this because of Scholia. Scholia is what prevents Wikidata from being a stamp collection, Scholia is what makes a subject come alive, it is what brings data together, makes it digestible and gives it relevance.

Not so for subjects relating to Wikimedia apparently for contractual reasons. There are several strategies to overcome this. But first let us decide what we are, what we do and why this matters.

Wikimedia is a publisher of scientific papers; currently there are three and in order to raise the impact of the papers it publishes, they have to gain visibility. To do this we can associate with ORCID, and publish and certify all the details of papers to its authors. One of the things we do on a big scale, is re-publish data from ORCID. They have a program whereby they can sync their information with ours.. They collaborate with Crossref and so could we. When we do, we make Open Science much more visible.

Dear Katherine, what we have shown is that we can and do care about publications, about citations. We care about science. The least we want is our own research to be presented the best we can. In order to achieve this we have to consider the unintended impact of a provision in a labour contract and overcome this self inflicted barricade.
Thanks,
       GerardM

Tuesday, December 18, 2018

#Wikidata and the papers of Professor Wiesje van der Flier

Professor van der Flier has an ORCID identifier. She works at the Neurology/ Alzheimer center of the VU University Medical Center.

Mrs van der Flier has in Iris E. Sommer a co-author. We know that they have at least one co-autor in Edwin van Dellen. There may be more and we will certainly know for those co-authors that are as open about their work. Professor Sommer was the initial interest because she is a member of "de Jonge Akademie".

We will know because they have an ORCID identifier. At Wikidata it serves two vital functions; it helps with disambiguation, a job was ran for all people with the surname "Li"... Given that ORCID allows people to share their information publicly, it allows us to import the publications of authors and identify their equally open co-authors.

The Scholia page for Professor van der Flier knew 31 people who were certainly knew to me. They are being processed and chances are that at the end of it Mrs van der Flier will know more co-authors, more papers and her representation in Wikidata will be more complete.
Thanks,
      GerardM

Yes, it will only know the co-authors that are open about their work but, that is only FAIR.

Sunday, December 09, 2018

#Science; I can read

The basis for what Wikipedia articles offers are its sources. Those sources can be anything and when we want to know the veracity of what we read, the sources have to be available. Not only that, we rely on those sources to be consistent and we rely on those sources to be readable.

When sources are on the web, the Internet Archive will have iterations of a source available in its Wayback machine. It ensures that sources remain available and thereby much of the integrity of Wikipedia is maintained.

For scientific sources we are unlucky. Reading a scientific paper can set you back $45,- and it only allows you to read that paper for a day.. In effect all such papers cannot be read; we "have to" trust them and there are plenty of papers that are extremely problematic and also expensive to read.

Many papers are increasingly FAIR. They are Findable, Accessible, Interoperable and Reusable. The best first line partners we have are again the Internet Archive and ORCiD. Organisations like the Biodiversity Heritage Library store scientific papers at the IA thereby making them available for as long as the IA exists. ORCiD is where living scientists identify themselves and if they so choose, the publications they (co-)authored. It makes them and/or their papers findable. The papers typically include a DOI making them accessible. After that it is anyone's guess if you can actually read them.

Scientists that are open about their work may find that they and their work found its way into Wikidata. For Karsten Suhre this was done; his scientific work is represented in his Scholia and many of his co-authors have been automatically added from ORCiD and have been processed as well. His co-authors that are not as open are largely missing but that is only Fair; I do not volunteer to promote them.

What Wikidata has is not representative of all of science but it increasingly represents the science that is open access, the science that I can read, that you can read that is for all of us there to read. The science that deserves to be used as sources in Wikipedia. We can read.
Thanks,
      GerardM

Thursday, November 15, 2018

Bringing more #science to @Wikidata

Slowly but surely more scientific papers and their authors find their way into Wikidata. Particularly when scientists have staked their claims in ORCiD, adding is easy and obvious.

It is easy because in ORCiD every author, paper, organisation et al have their own unique identifiers. So when you add a paper, all authors who claimed to be author are already linked.

Earlier today, I added papers and co-authors for Jaume Piera. As a consequence Laura Recasens was added today and as you can see in the illustration of her co-authors, several new authors popped up as a consequence.

To do this I use a combination of tools. Reasonator is my preferred tool to display data; for scientists it tells me if he or she is known to be an author. When there are, Scholia presents the scholarly author information. Of particular relevance to me is the co-author presentation. For co-authors shown in white, no gender is given in Wikidata and when the name is an initial and a surname, I will look up the ORCiD information to find a full name. Typically that is how people are known in ORCiD.

I use the SourceMD tool for two purposes; "creating and amending papers for authors" and to "add metadata from ORCiD authors to Wikidata". It is processed in a batch job, I run one job for up to 15 authors at a time and it takes forever to run.

Other people run other jobs, a particular hat tip to Daniel Mietchen who makes sure that recent publications find their way into Wikidata and finds many other reasons to improve on what we have. All this would not be possible without the many tools by Magnus and for Scholia I do thank Finn Ã…rup Nielsen thanks to this evolving presentation, science as a process comes alive.

There is more to do; the Wikipedia citation are in a separate database and much of its data may be found in Wikidata.. Who will merge them. Publications do cite other publications, it is a field I am not really interested in.. They are added so there must be a tool.

When you are interested in a particular scientist, a particular paper.. Just use the tools and slowly but surely we all make Wikidata a great tool to represent science fact.
Thanks,
      GerardM

Saturday, November 10, 2018

More #impact for your #science is in being a #source at @wikipedia

In a study about how students research a new subject it was found that they read the Wikipedia article first. Then they move to its sources and from there it takes off.

In order to have an impact you, as a scientist, wants to be their first getting the attention of your work. There are a few tips.
  • Make sure that you and your work are known. First make your work known at ORCiD. From there it gets into Wikidata
  • PS check out the Scholia presentation of you and your scientific work.. (example)
  • Make sure that your work can be read. Wikipedia actively seeks free reads using the OAbot.
  • Do not think that current practices of your field will benefit new scientists in the future. Many fields are not well represented at ORCiD
For your information. There is a database with the sources used in Wikipedia. The only thing lacking is that this database still needs to be integrated in Wikidata for it to gain a real impact.
Thanks,
      GerardM

Saturday, October 27, 2018

#Library #Science - Prof Dr Frank Huysmans

Mr Huysman's works at the Universiteit Amsterdam. He teaches "Library sciences" and as is usual for a scientist, he has a fair share of publications to his name.

The problem is that this field of science is not well represented in Wikidata. There were no publications to his name. Importing them from ORCiD proved problematic; only four were added out of the 22 known there. Working from what was known, it was possible to add co-authors and enrich those, seek out their co-authors and enrich them as well. The result is the current 40 publications to Mr Huysman's name.

Mr Huysman has both a Twitter and an ORCiD account. Everybody who does, in Wikidata, will have his or her profile in Wikidata updated thanks to a job that is running by Daniel Mietchen. They are the ones who publicly promote their science and in this way they gain some additional credibility.

NB when you have an ORCiD and twitter, tweet #IcanHazWikidata and you will get your Qid.

When you care about your science, do maintain your ORCiD profile because it will make your papers, your co-authors and the organisation for more visible in Wikidata.. Your #Scholia profile will get better and better and chances of being quoted in Wikipedia improve.
Thanks,
     GerardM

Monday, October 22, 2018

#Science - Ladies you work together

Yesterday I singled out a Paola Giardina because she was a co-author of someone who had SO many co-authors, I could not manage the information that was in there. Yesterday Paola had a large number of co-authors that were white (no gender info). Today there are even more present.

One thing is pretty obvious in what I see: women are more likely to work with women than men. When you want to analyse this, it is important to know the data this is based on. At this time 31% of the people with an ORCiD identifier are female. When you consider probability, it is likely that some 31% of people who have not been associated yet with a gender will be female as well.

In many universities the percentage of women studying is more than 50%. All of them get involved in research. All students are involved in the production of papers and all of them are entitled to their ORCiD and to their Wikidata identifier.

So when we want to express the notability of women in modern science, all we have to do is ask any and all scientists to make their publication details part of the open record. Slowly but surely, it will become obvious who and where the best science is produced and who collaborates with whom.
Thanks,
     GerardM

Saturday, October 20, 2018

#Accepting science; the solution is in the reading not the publishing

The most important thing religion has over science? Its papers can be read. Sources like the Bible, he Quran can be read for free. You can get *your* copy from many true believers. A copy is in your library. With science the papers that can prove to you that goldfish should be classified as endangered are behind a paywall. It is only your common sense that might say: "Hey, wait a minute.."

When Wikipedia insists on its sources, they are only functional when these sources can actually be read. This is why the Internet Archive plays such a vital role in maintaining the validity of stated facts.

Some scientists think that "the public" cannot read scientific papers. They forget that even for scientists a paper that cannot be read is a paper that does not exist in their contemplations. The public does read scientific papers. The Cochrane crowd for instance reads papers and checks particular premises for validity.. We know that scientific research of coronary disease was biased for males and as a consequence women still die. A bias like that is what they look for, it is why they reject many papers because they are basically *not* valid.

There is a lot to do about what scientific publishing should be. How it should be funded.. The base line is that when a publication is not available for anyone to read, the facts do not matter. Why believe vaccines are safe when the publications that prove it are behind a paywall?
Thanks,
       GerardM

Friday, October 19, 2018

#Wikidata - the missing #Elsevier papers

It started with a Twitter tweet.. "There is also a professor Elsevier". A search found that Professor Cornelis J. Elsevier works at the "Universiteit of Amsterdam". He did not exist at Wikidata and there was only one paper to be found for him.

Adding this one paper was done with the "Resolve Authors" tool. The Scholia tool for Mr Elsevier showed a few co-authors and in addition to this several "missing co-authors" could be found.

In order to show more papers for Mr Elsevier, more papers needed to be imported into Wikidata. This can be done for authors with an ORCiD identifier, particularly the ones with no known gender. So far they did not get much TLC. Just running the "SourceMD tool" for them will add additional papers and associate other authors to these papers as well.

This is an iterative process and I focused for no particular reason on Mrs Barbara Milani. Processing her co-authors meant that more co-authors came out of the woodwork. At this time, 13 new authors with an ORCiD identifier popped up. Once they are processed more papers will be known to Wikidata and given their relation to Mrs Milani a reasonable chance that these papers link to Mr Elsevier as well.

At this time Mr Elsevier is known to have 7 publications.
Thanks,
        GerardM

Sunday, October 14, 2018

#Wikidata - the #heart of women differs from the heart of men

The assumption that the heart of women and the heart of men are the same proved to be lethal. The "Hartstichting" is a Dutch charity that raises funds to combat heart disease. One of its studies is done by professor Hester den Ruijter of the Utrecht Medical Centre. Her study aims to map those differences and it is part of an effort to provide equal quality medical support for heart matters for both genders.

As a scientist, Mrs den Ruijter was involved in the production of many scholarly papers with many co-authors and this is best presented by Scholia. Yesterday Mrs den Ruijter was only known to Wikidata through her papers. Today she has her own item, the papers have been associated with her and so have been many of her co-authors. Many other authors have their own item who are associated with the research that indicates how the heart and its diseases differs between the genders and differs based on ethnic background.

It is vital to recognise these differences, survival relies on it.
Thanks,
      GerardM

Sunday, October 07, 2018

#WikiCite - Thank you #Orcid ! - #IcanHazWikidata

The question "I have an ORCiD profile, how do I get it in Wikidata" was asked on Twitter. Using Magnus's tool public information was imported and as a result information can be shown in Scholia.

Paolo Cignoni made a request using the #IcanHazWikidata hash tag and his papers were imported and it shows nicely in Scholia. It includes several of his co-authors, for the ones in white we have no indication for their gender in Wikidata. That is easy to fix.

There are probably a lot of co-authors missing.. One way of finding the missing co-authors is by adding "/missing" to the Scholia link. You can check for an ORCiD identifier and add a found identifier. You identify the papers already known to Wikidata and they are attributed to the co-author or, to a citing author.. I added a John W Goodby to make the picture more complete. It is easy and mostly obvious what to do.

What makes all this possible? Open data and a bit of effort.. As you can see in the later picture, just running Magnus's tool for a few co-authors changes the outlook considerably.

Are you a scholar and do you want to see your initial Scholia information? Just add your Ordid ID in a tweet with the #IcanHazWikidata hash tag.
Thanks,
      GerardM

Wednesday, October 03, 2018

#Wikimedia - Relevance of #science - Kate Ricke

A lot of soul searching happened to determine why Wikipedia failed to notice Donna Strickland only once she received the Nobel Prize.. What is more astounding is that Wikidata failed to include her.. No Scholia information for her and her research. What we have at this is likely to be a subset of the "Stricklands papers".

We do not know who will be seen as a scientist of similar relevance but we do know that a lot of rubbish is floating around.. it is called fake science, fake news and countering this is where big organisations like Google and Facebook rely on the information in Wikipedia.

So Mrs Kate Ricke is another scientist that did not get Wikipedia attention so far. Mrs Ricke tweeted about her paper Country-level social cost of carbon. It and the papers produced by her and her co-authors are quite potent.

When you learn about a paper like this, you can add it and its authors to Wikidata. When Orcid has information about other papers, you can import these papers as well building on the web of science about of one of the most important subjects of our time. In addition co-authors of these other papers can be included as well as the authors citing these papers.

When relevance is given to the science of a subject like climate science, it becomes possible to contrast it with what some politicians want us to believe.
Thanks,
       GerardM

Sunday, September 30, 2018

#Wikimedia supported scientific papers supported by #Scholia

There are four scientific journal published by the Wikimedia Foundation they are:
According to an interview, they offer articles available under a free license at no publication cost. With a platform for publications the next thing is to gain notability for the journal and for its authors.

One way to assess the value of these papers is by checking what Scholia has to say about these papers:
Part of the Scholia information are the links to the authors; not only who has been most prolific as an author but also who has been cited the most. There is one caveat; the author needs a Wikidata item and the more complete the information, the more both the journals and the authors gain in notability..

PS Ladies, the ratio men and females is not really what I would expect for a Wiki journal..
Thanks,
      GerardM

[1] The WikJournal of Humanities is under development; its first publication is in the future at this time.

#WikiCite: Edith Abbott; THE REAL JAIL PROBLEM

According to #Wikipedia, Edith Abbott was a pioneer in the profession of social work with an educational background in economics.She published, books and scholarly articles, and "The Real Jail Problem" was singled out in the article for special mention.

There is a link in the article, and I am happy to report that thanks to the Internet Archive, it and some 500.000 more scholarly articles may be found there. They are part of the early JSTOR papers and they are now freely available. It is however not the publication by Mrs Abbott, it is by a Robert H. Gault.

Having scientific papers available is wonderful. It is important and it is unlikely they will get lost. However, there is a difference between being lost and being findable. Mr Gault is notable, there was no mention of him in Wikidata, one of his publications was.. Mr Gault is findable when you look him up in VIAF.

However, the article you want to read, a pamphlet according to Mr Gault is where? A book with the same subject can be found at Open Library. The subject of the book is as potent as it was in the day. The arguments are not dissimilar.

Arguably, when you want to cite your sources, they first have to be knowable, then findable and finally readable. Wikidata is making more and more of a difference. When we collaborate with the Internet Archive, we can make those JSTOR papers findable as well. Slowly but surely all these sources becomes findable. Its authors become notable and we will free us from the curse of only reading the latest research.
Thanks,
       GerardM

Tuesday, September 25, 2018

#Wikidata: Adding credibile info to #Twitter - Margaret Stanley

Twitter, Facebook, organisations like them have this credibility problem. Too many Joes Dicks and dirty Harries poison the well that is the information they provide.

Now meet Margaret Stanley, she is a scientist and, she tweets. Her Twitter name is included in Wikidata and it is highly likely that even though her tweets are her own and, do not indicate the opinion of the institution she works for, what she tweets is credible and well considered.

Margaret is not the only scientist that tweets, there are many of them. More and more of them are included in Wikidata, including their publications, including their twitter handle. One of these scientists, actively encourages female scientists to speak out, seek a platform to encourage women to find their place in science. They twitter, they write Wikipedia articles and they are very much relevant scientists, much of their relevance shines through in their tweets.

Dear Twitter, when scientists have a profile in Wikidata, they personally make a statement. It is theirs. They are as human as everyone else but they are not, as a group, foolish enough to tweet balderdash, nonsense or other stuff you should frown upon.
Thanks,
      GerardM

Tuesday, September 18, 2018

#Wikicite: #Research? Eat your own dogfood!

Another piece of Wiki research has been published and you would not know when you look for it in the Wikiverse. The people who are big in Wiki research run this project called Wikicite. It has multiple aspects; publications and there authors are entered in Wikidata. Blogs are written how wonderful the various aspects are of what is on offer in the growing quality and quantity of data and how well an author may be presented in Scholia.

This is all well and good and indeed there is plenty to cheer about like the documentation about the Zika fever that is included but when a subject that is key to the Wikimedia researchers is not as well represented it will never be good enough.

There is a reason why using your own is so relevant; it shows you where your model fails you. WAll kinds of everything like conference speakers are included, many more female scientists are represented as a result of the indomitable Jess Wade but when new research by Wikimedians, professional Wikimedians, are not included, the effort is not sincere; the people that do go to conferences do not learn from the daily practice and that makes Wikicite stale and mostly academic.
Thanks,
     GerardM

Monday, September 03, 2018

#Wikidata - Presenting Mr Tanvir Hussain

Mr Hussain, a professor at Nottingham University mentioned on Twitter what he would put on his business card. Many of his publications were on Wikidata and after some additional information from Orcid, Mr Hussain looks really good when viewed with Scholia.

He asked me if we could have the Scholia information with a QR code like we do in Reasonator for people with a Wikipedia article.

Having a QR link to information like Scholia would look really good on a business card.. It would also stimulate the colleagues of Mr Hussain to get this well represented.
Thanks,
     GerardM

Tuesday, August 28, 2018

#WikiCite - A #Cochrane Pocketbook #Pregnancy and #Childbirth

In Wikidata many, many publications representing scientific publications from everywhere have found an entry. There are literally millions of publications, in essence it is an increasingly comprehensive stamp collection in need for some structure. Without structure there is no use for it. 

There are many publications about pregnancy and childbirth. When you check out the links, you will find that currently publications about pregnancy are very much dominated with the Zika virus and for childbirth Mr G. Justus Hofmeyr is mentioned as one of the authors important for the subject.

Mr Homeyr is mentioned only because the Cochrane Pocketbook "Pregnancy and Childbirth" was included by hand. This book is quite significant because it represents the best evidence for doctors and other medical professionals who provide maternity care.

As more structure is given to publications, it becomes more useful; publications will gain authors all known individually to Wikidata. The citations of publications will be mapped and publications of Cochrane will bring out what publications truly add to the sum of all knowledge.
Thanks,
       GerardM

Monday, August 20, 2018

#Wikidata - Never heard of "R Andrew Moore"?

When quality in scientific papers, papers that are to be used as sources in Wikipedia, is important, it is relevant to include the papers published by the Cochrane Database of Systematic Reviews.

For one author, R Andrew Moore, I invoked the QuickStatementsBot and he added Mr Moore, added some 131 links to publications. It was a subset of what could be done but hey I already felt I was living dangerously.

When you then have Scholia look at Mr Moore for missing information .. there is a lot. But there is so much information.  In the diagram you see his co-author graph, and you will also find that Mr Moore published 123 times for Cochrane.

There are many publications and arguably we may want them all when we are to share the sum of all knowledge. Not all publications are equal and  Cochrane is special because it reviews what is published elsewhere; their bottom line is not commercial, their motto is Trusted evidence. Informed decisions. Better health. It is what aligns best with what we aim to do.

As I learn how to add authors who published for Cochrane, I will do maybe one a day. When other people take an interest, slowly but surely meta data on research gains relevance in Wikidata and with a little help of our Wikipedia friends we will provide better information.
Thanks,
      GerardM

Saturday, August 18, 2018

#Wikidata - Speakers at the #Cochrane Colloquium 2018

All the speakers announced for the Cochrane Colloquium 2018 have a presence on #Wikidata. Cochrane is a vital organisation; it debunks much of fake science by researching published papers. It asks for volunteers to help with this effort and its mission is to "promote evidence-informed health decision making by producing high-quality, relevant, accessible evidence".

Cochrane and the Wikimedia community collaborate, there is a Wikipedian in Residence. There have been editathons where articles were updated based on evidence (not sources) and consequently there is ample room for collaboration at the Wikidata level as well.

All speakers have been added to Wikidata and Mrs Sue Ziebland a professor at Oxford University, was new. This was not the case for many of her papers; they identified one author as "Sue Ziebland". Daniel Mietchen was so gracious to link the papers to the person and thanks to the Scholia tool Mrs Ziebland gets a lot of depth.

So what could a Wikidata / Cochrane collaboration look like? With a positive spin, all the positive authors according to Cochrane get a Wikidata item and, like Daniel did for Mrs Ziebland, the publications are linked to the person. When Cochrane provides links to its database, we can even see why a specific paper is so relevant. This will help Wikipedia editors because evidence-informed health decisions can be made and determine disputes in articles.
Thanks,
        GerardM

Sunday, August 12, 2018

#Knowledge - three types of knowledge and why "academic" is only one and overrated

There are three types of knowledge; they are academic, professional and knowledge from experience. The scheme to the right was published by Jaap van der Stel. He works in the field of psychiatry and is known for his work on addiction in combination with the use of peers in the recovery from addiction.

In the Wikimedia world, we insist on the primacy of academic knowledge and up to a point it serves us well. Operationally it means that much of the studies are done outside of the WMF, they may point out whatever but they hardly ever make an operational difference. When the internal WMF researchers study a subject, they are typically directed to study particular phenomena and it may point to operational issues. Issues that are either addressed by the WMF itself or adopted by the community.

When scientists make a compilation of all the sources in all the Wikipedias, it is academic work when the result is static. It may indicate what sources are used multiple times but it does not help any editor weed out sources that are biased or false. Magnus started work on a tool that knows about all the sources in two Wikipedia and Wikispecies.  It is updated in real time and that  gives it valid operational credentials.

I know from experience that there are issues with source information as we have it in Wikidata. We cannot invalidate sources by reference. We are only strong in the biomedical field and adding new information is not at all user friendly.

Now this user experience does not get much of a priority for valid operational reasons but the effect is that Wikidata is only useful for the geeks. Its lack of usability prevents its data to be used on Wikipedias in the "other" languages. It is where there is little or no academic nor operational interest.
Thanks,
      GerardM

Saturday, August 11, 2018

#GenderGap - The Gineta Sagan Award (and others)

The Ginetta Sagan award is conferred by Amnesty International USA. It is an annual award, the last recipient according to Wikidata when I looked at it received it in 2014, English Wikipedia has the award as part of the article on Ginetta Sagan and has information including 2017 (when you read the texts, you will find how notable these people are and, by inference the people without an article).

Arguably, there is a lack of balance between the number of men and the number of women having an article in any Wikipedia. This is known as the "gender gap" and the "women in red" project works to great effect to improve that balance. There is no lack of fine notable ladies who have no article.

I am really happy to present two queries. The first query shows women who won an award with no article at all (2502 results). The second shows women who won an award with no article in the English language (29083 results).

Let these women be an inspiration to you.
Thanks,
       GerardM

Monday, August 06, 2018

#MADinAmerica - cause and effect

MAD in America is an organisation about mental health, particularly in America. Their take is that there is a lot that can be improved. The part that I am mostly interested in is that they highlight the science that tells you how the science behind many mental health practices fails scrutiny.

One publication they recently highlighted is about brain abnormalities by people with schizophrenia. Current wisdom has it that "cortical thickness and surface area abnormalities in schizophrenia" is indicative of schizophrenia. This paper compares people with schizophrenia who were medicated and people who were not medicated. The research shows that these differences are due to the medication.

Adding a paper like this in Wikidata is easy. Making it stand out for its results is not. The paper probably indicates previous research that it debunks but how do you model that. When papers like this are to be used as sources, how do you ensure that it is even considered?
Thanks,
      GerardM

NB the first author is employed by the University of California, Irvine

Sunday, August 05, 2018

#Citations - "Verlorene Siegen" and #Wikipedia

The publication The Battle for Wikipedia: The New Age of ‘Lost Victories’? writes about debunked knowledge but used as sources in Wikipedia. Lost Victories is a book by Erich von Manstein, a German military officer convicted at the Nuremberg trials. He served as a witness and there is strong evidence that he perjured himself. He was sentenced to eighteen years in prison

This publication is not only of academic interest. In this day and age where fake facts and science are pervasive, it is a reminder that Wikipedia is a battle ground where debunked sources are used to prove a non-neutral point of view.

One of the objectives of the Wikimedia Foundation is to combat fake facts and use make citations operational as a tool. The main trust will be by adding sources to Wikidata. "Verlorene Siegen" obviously was present but even though there is a large body of work debunking this book, there was nothing to refer either to Mr von Manstein or his book in a critical way.

It was easy enough to add a few individual sources but it takes time. For analysis of sources used in Wikipedia there are dumps containing all the citations of all Wikipedias and now Magnus has started on a tool that initially includes real time sources for the German, English Wikipedia and Wikispecies. Of these publications 36% are linked to Wikidata and this provides a great start but it will take more. We need to know what papers debunked what knowledge. We need to know what papers a Retraction Watch is critical of, or what the relevance of a paper is according to the Cochrane Database of Systematic Reviews because that is how their facts are operationalised. We need to know because that is one way to debunk fake facts.
Thanks,
       GerardM

Saturday, August 04, 2018

#Wikidata - User versus bot updates and #Scholia

These are the aggregated subjects that are associated with all the papers for the winners of the Fields Medal. Given that there are some 60 award winners for the most prestigious award in the field of mathematics, this is not a representative reflection. That is not a problem, that is an opportunity.

I added one paper, "Singularities of linear systems and boundedness of Fano varieties". Given the title, I added "Fano variety" and "Linear system" as subjects. This made no difference in the Scholia tool and after some five minutes I asked what was happening. I was told that it takes a large interval before the data in the Toolserver get updated.

Typically, information about papers are added by bot. Not so much for mathematics but still. Mr Birkar for instance has only two papers in Wikidata at this time and for the other paper no subjects are given. When you add data by hand, instant gratification or instant visibility is important as it is a potent motivator.

The best reflection of work done in Wikidata is not given by Wikidata itself. It is either by tools like Scholia or Reasonator or it is by query. When query does give instant gratification, it has much of its potency because of the instant gratification.

Tools have one important benefit over query; it provides a standard layout for the information. Queries are potent and many people contributing content to Wikidata use it in tools like Petscan. But in reality, the typical difference between one query and the next are only in the qualifiers.

At this time the best user experience is given by tools. It often suffers from a time lag and this is of little relevance to bots. For humans though it is different.
Thanks,
       GerardM