Showing posts with label Biodiversity Heritage Library. Show all posts
Showing posts with label Biodiversity Heritage Library. Show all posts

Sunday, August 11, 2019

How to value open data and why Wikidata won't go stale

The data in Wikidata is data everyone knows or could know. A lot of awful things could be said of its content and quality and all of it misses one important point. It is being used, its use is increasing, it is increasingly used by Wikipedias and that provides an incentive to maintain the data.

What Wikipedia indicates is that most data is stable, not stale. A date of birth, a place of birth so much remains the same. When we bury data in text, it is always a challenge to get the data out. When we bury data in Wikidata it just takes a query to bring it back to life. Who was a member of multiple "National Young Academies, Similar Bodies and YS Networks" for instance; you do not find it in the texts of those organisations but you will increasingly find it in Wikidata. Once the data is in there, it is stable and available for query.

As GLAMS make their content available under a free license, their collections gain relevance as the collection gains an audience. Just consider that only a small part is available to the public in the GLAM itself and on Commons it is there for all to find. Commons is being wikidatified and those collections become available in any language gaining additional relevance in the process.

The best example is what the Biodiversity Heritage Library does. It is instrumental in the digitisation of books, it makes them publicly available and gains the collections they are from an audience. Volunteers prove themselves in this process and both professionals and the wider world benefit. From a data perspective the data is new because only now available.

When a publisher mocks Open data, it is self serving. It is in their interest that data is inaccessible, only there for those who pay. There are plenty of examples of great data initiatives that went to ground and obviously when the data does not pay the rent, publishers will pull the plug. It is different for the data at Wikidata. It is managed by an organisation that has as its motto "share in the sum of all knowledge". The audience the WMF has makes it a world top ten website, it is not for sale and it is not going anywhere. As long as there are people like me who care about the availability of information, the data at Wikidata may go stale in places waiting for another volunteer to pick up the slack.
Thanks,
      GerardM

Sunday, December 09, 2018

#Science; I can read

The basis for what Wikipedia articles offers are its sources. Those sources can be anything and when we want to know the veracity of what we read, the sources have to be available. Not only that, we rely on those sources to be consistent and we rely on those sources to be readable.

When sources are on the web, the Internet Archive will have iterations of a source available in its Wayback machine. It ensures that sources remain available and thereby much of the integrity of Wikipedia is maintained.

For scientific sources we are unlucky. Reading a scientific paper can set you back $45,- and it only allows you to read that paper for a day.. In effect all such papers cannot be read; we "have to" trust them and there are plenty of papers that are extremely problematic and also expensive to read.

Many papers are increasingly FAIR. They are Findable, Accessible, Interoperable and Reusable. The best first line partners we have are again the Internet Archive and ORCiD. Organisations like the Biodiversity Heritage Library store scientific papers at the IA thereby making them available for as long as the IA exists. ORCiD is where living scientists identify themselves and if they so choose, the publications they (co-)authored. It makes them and/or their papers findable. The papers typically include a DOI making them accessible. After that it is anyone's guess if you can actually read them.

Scientists that are open about their work may find that they and their work found its way into Wikidata. For Karsten Suhre this was done; his scientific work is represented in his Scholia and many of his co-authors have been automatically added from ORCiD and have been processed as well. His co-authors that are not as open are largely missing but that is only Fair; I do not volunteer to promote them.

What Wikidata has is not representative of all of science but it increasingly represents the science that is open access, the science that I can read, that you can read that is for all of us there to read. The science that deserves to be used as sources in Wikipedia. We can read.
Thanks,
      GerardM

Monday, December 18, 2017

Francesco Redi and the #BHL - A purpose for #Wikisource

Mr Redi is of a stature that his statue is in the Uffizi Gallery. His books are available in the Internet Archive, thanks to the Biodiversity Heritage Library.  Wikisourcers waved their magic on several of his books and the result is a superior output for for instance "Esperienze intorno alla generazione degl'insetti".

Mr Redi has four books who received the Wikisource treatment.. and then what? In response to a tweet, I was told of the existence of these books in Wikisource. I checked them on Wikidata and added Mr Redi as the author. There is nothing to indicate in Wikisource where the book came from (the BHL provided them with a DOI).

In a tweet, the BHL indicated that they are interested in books that received the Wikisource treatment. So lets consider where we are:
  • Wikisource has many great books transcribed and available as an ebook
  • It is not known outside of Wikisource what books are available in what quality and where they came from
  • We could have this information in Wikidata. It will give a clue what is available; we can query for the books when they are in Wikidata
  • What is the purpose of Wikisource if it is not for people to read all these fine books?
Thanks,
       GerardM

Thursday, December 14, 2017

A purposeful #strategy for #Wikidata

A strategy for Wikidata? Obvious, it is all about having a purpose. It is not about policies, it is not about what we need or expect of others but it is about the purpose you, I and others have for us to collaborate on in an inclusive Wiki and data project.

The implication of making the purposes of our community rule supreme are huge. Purpose like so many other things can be measured. When people have a purpose for Wikidata and actually use it, their need for quality is self evident. They will invest their time and effort in fulfilling their purpose. The one question is how to fit in the many purposes that exist for Wikidata.

Take for instance the objective of Lsjbot for a rich Wikipedia in the Cebuano language. He uses data from an external database to create articles. Data from these articles are imported later through the Cebuano Wikipedia in Wikidata. This is seen by some as controversial because of the need to integrate data that often already exists. The purpose is obvious; rich information in the Cebuano language. The solution is obvious as well; let Lsjbot use the data at Wikidata to generate the information for the Cebuano Wikipedia. GeoNames is happy to collaborate with us on this, so when we care to collaborate and welcome its data at the front door, we can mix'n'match the data into Wikidata, curate the data where necessary and share improved quality widely, not only on the ceb.wp.

The Biodiversity Heritage Library Consortium is working extremely hard to expose their work to the general public. Over a million illustration found their way to Flickr. Fae imported many of these to Commons and most if not all the associated publications can be read on the Internet Archive or on its website. Their content is awesome, check for instance their Twitter account. We can import all the BHL books in Wikidata, we are importing all associated authors using Mix'n'Match. The images are in Commons but how is this brought together? How do we add value for the BHL and as important, for our shared public?

The Internet Archive is a Wikimedia partner. It provides essential services for us with its "Wayback machine". It is how we can still refer to references that used to be online. One other venture of the Internet Archive is its Open Library.  What we already do for the Open Library is linking their authors and by inference books to the libraries of the world through VIAF. We could share this information with the Wikipedias so that its readers may find books they can read. (Talk about sharing the sum of all knowledge).

Both the IA and the BHL want people to read. They (also) provide scientific publications that may be read to prove the points Wikipedia authors make in articles. Both can be big players strengthening the value of citations in WikiCite. At this time its strength is particularly in the biomedical field and it is already attracting bright people to Wikidata. As data from other fields finds its way, people like Egon and Siobhan will find their way. This will make Wikidata even more inclusive.

To make this future work, to become more inclusive, we should trust people more particularly when they indicate why they use Wikidata. The Black Lunch Table is a great example. The description at Wikidata says: "visual artists of the African diaspora initiative that includes Wikipedia editathons and outreach". One way of knowing how effective this initiative is is the history page of its listeria list. It shows a steady growth of information added. When you analyse it further you find artists added and selected for new editathons. Truly a great example of Wikidata having a purpose.

A strategy based on purpose, is a strategy based on trust. Not blind trust, but the kind of trust where it is seen that people are committed to improve both quantity, quality and usefulness of the data they identify with.
Thanks,
      GerardM

Tuesday, November 28, 2017

#Wikidata - Disambiguating for the Biodiversity Heritage Library

Tatiana Carneiro is an entymologist. Her work is known at the Biodiversity Heritage Library. When you check the "authors page", there are two other identifiers known and, for "Tatiana R. Carneiro" the same two identifiers are shown as well.

When you google for Mrs Carneiro all kinds of information may be found but you do not want to do this for all the 177,271 BHL authors that are waiting in Mix'n'Match. It is no fun and only a few people take up a task like this.

So the question is; how do we make it more rewarding and how do we bring the many Brazilian papers to Wikidata as well. What is it that there is to achieve and how does it benefit all the people reading Wikimedia content.

For readers of our content, there is little merit in the fact that all these authors published papers. Many of them have been published with a DOI and, many of these papers are freely available to read. For them the papers are important. So contrary to a more normal database approach it is not the authors we should concentrate on but it is their publications. In addition to this, the BHL actively promotes the use of illustrations and publish them on Flickr. Thanks to the fine work of people like Fae these illustrations end up on Commons as well. It will be a challenge to link them to all this metadata..

There are millions of illustrations, there are far fewer publications and many authors are known for not one but multiple publications. To complicate it even further, an illustration has an illustrator and many publications are exclusively found in archives. Many publishers are no longer active and all this information is or may be considered relevant.

So what to do; first import all the publications that are freely readable. The publications with a DOI and include the author information as "author name string". When an author is known to Wikidata, we can always add the author information as well. The benefit of this approach? People can read now.

To make it interesting we can run a bot using the APIs of the BHL. We add missing books for authors and add the authors to the books where this information is missing. Running this regularly will make it interesting for anyone interested in the work of the BHL. But most importantly, people can read now.
Thanks,
      GerardM

Saturday, August 05, 2017

#Wikidata - Harriet Martineau and some social opportunities

When you do not already know about Mrs Martineau, do read one of the many Wikipedia articles, she is considered to be the first female sociologist and introduced many subjects into sociology that were up to that time not considered.

The picture is a crop of a painting at the National Portrait Gallery by Richard Evans. The picture is known at Wikidata, at Commons the Creator template is missing.

At the Biodiversity Heritage Library Mrs Martineau was know for her book a complete guide to the English lakes. It was the only book known for her at Open Library.  Given the relevance of Mrs Martineau this was strange and sure enough she was known as "Martineau, Harriet" and changing the link to the book was easily done.

At Wikidata meanwhile, there was a hidden link to Mrs Martineau to Open Library thanks to all the good work of the Freebase volunteers. Approving the change was obvious.

At Wikidata there is now a link to both VIAF, to the BHL, to OL for Mrs Martineau and to over 20 more sources. The BHL has links to both Open Library and VIAF. When the links differ, it becomes obvious where work needs to be done.

The result is a better service for all the people who make use of any or all of these resources. We truly should collaborate and strengthen our partners, the partners we share data with.
Thanks,
      GerardM

#Standards - the International Plant Names Index

#IPNI is a collaborative project between three august bodies in the taxonomy of plants. They are the Royal Botanic Gardens, Kew, the Harvard University Herbaria, and the Australian National Herbarium.

There are three areas where IPNI sets the standards: plants, authors and publications. The objective is to disambiguate any taxonomic reference to a plant in scientific literature to the correct taxon given the taxon name, its author information, publication information and date.

IPNI publishes several graphs indicating the success of their work. I have been involved in this work as a consequence of a database project I did for my father who loved his cacti and succulents.

One example of what information IPNI provides can be found in this page for the "genus" Echninocactus. In my understanding, the correct full taxonomic name is: "Echinocactus Link & Otto Verh. Vereins Beford. Gartenbaues Konigl. Preuss. Staaten 3: 420. 1827". It has all the required information, it has type information, it has links all as you would expect of a standard like this.

To appreciate the work of IPNI; in stead of "Link & Otto", there may have been: "Link and Otto" or "Link et Otto" or ... obviously the information for the publication is easily made into a different abbreviation.

Wikidata included only a subset of the full taxon information. It is easy enough to understand why; Wikipedia only needed the most current one. It is an easy model; works relatively well and it breaks in the corner cases. With the development of WikiCite there is a great and possibly easy opportunity to expand on the current work given the expanding collaboration with botanical partners like the Biodiversity Heritage Library.
Thanks,
      GerardM