With currently over 12,000,000 "items" registered in Wikidata, most of them registered with bots there is a monumental task waiting to be undertaken. It is adding sources to all this information stated as facts.
In Wikipedia it is policy that facts stated in an article need to be supported by sources. It is also a matter of principle that Wikipedia itself can not be considered a source itself. Stating that something is true because Wikipedia says so is good for more than a smile. Oh and, there is not one Wikipedia, there are over 280 Wikipedias; enough reasons to snicker.
One of the things people are taught in Islamic schools is that they should rely on the original sources. When something of a religious nature is stated, it should be supported by what can be read in the original sources of the Islamic faith. In Wikidata many people and subjects that have to do with Islam have found their way as a fact that is not sourced. They include genealogical information like the one shown above or similar information about Q9458.
Having sourced information is important because some information present in Wikipedia is certainly wrong. Having incorrect information in Wikidata is even worse because it may present information used in 280 Wikipedias.
Bringing together people who know the relevant sources, who are willing to learn about Wikidata and edit its information is something you can do in a workshop. Organising a Wikidata workshop together with a mosque are two novelties; as far as I am aware there have been no workshops organised around Wikidata and, organising a Wiki workshop with a mosque is something I have not heard about either.
Thanks,
GerardM
Sunday, May 19, 2013
Tuesday, May 07, 2013
#Lua template wanted for use of #Wikidata in #Wikipedia
Denny said some magical words to me;
- you can convert a list of Wikipedia wiki links to links to Wikidata using Lua.
- you can check if there is an article on THIS Wikipedia
- you will show the Wikidata data when there is no article
He even said that it is not hard to do and provided this as a pointer.
I am preparing a Wikidata workshop and I would be REALLY pleased when this list was available doing all the song and dance mentioned above. There are so many other lists that could benefit from this as well.
I am convinced that such functionality will motivate people to write stubs and articles on subjects that are important to them. I am also convinced that it is a powerful incentive to create data that accomplishes things like this.
Thanks,
GerardM
Thursday, May 02, 2013
#Wikipedia lists could fall back to #Wikidata
I have been playing with Wikidata and it is really good fun. I find many uses for it and some of them have to be tweaked a bit to be even better. Take for instance a list of popes. There is a list with articles on the English Wikipedia for each of them. There are so many popes, that it is obvious that many Wikipedias do not have the list and certainly not articles to all of these popes.
Wikidata could come to the rescue. When a list is made up of values available in Wikidata and when the links to articles fall back to Wikidata, we are able to provide relevant information and, we have the perfect opportunity to suggest to our readers to write a stub or an article.
In effect such a list allows us to provide improved information by using the strength of Wikidata in any of the languages we support. When you think about such a list, it is not much different from an info-box.
Thanks,
GerardM
Wikidata could come to the rescue. When a list is made up of values available in Wikidata and when the links to articles fall back to Wikidata, we are able to provide relevant information and, we have the perfect opportunity to suggest to our readers to write a stub or an article.
In effect such a list allows us to provide improved information by using the strength of Wikidata in any of the languages we support. When you think about such a list, it is not much different from an info-box.
Thanks,
GerardM
Wednesday, May 01, 2013
#Wikipedia red links should link to #Wikidata
When a subject does not "merit" an article in Wikipedia, it does not follow that the same subject is not of value in Wikidata. Consider for instance the son of a famous person who died at the age of three. He completes the list of all the children of that person something that is definitely "a good thing" in Wikidata.
It is certainly true that many subjects that "merit" an article do not have an article. There may be an article in one language and not in another. Wikidata has the option to add a label for such an article as a place holder and while it has not been written it can show nicely red on a disambiguation page. The point here is that it is legitimate for lists and disambiguation pages to have red links. When such lists are completed in Wikidata they can easily be translated to other languages and provide basic information that may be of interest.
The red links in the list of Muhammeds are both kings of the Sayfawa dynasty. Sadly they are not even all the kings called Muhammed who are part of the Sayfawa dynasty with a red link. Just consider what would happen when disambiguation lists are presented from Wikidata; it makes it easier to start articles because relevant information may be available thanks to work done in another language.
When you consider the options, all "red links" could be known to Wikidata. As a result you can complete all lists without having to write articles and, you will deal with disambiguation issues sooner rather than later and, why have wikilinks when integrity is better maintained in Wikidata anyway?
Thanks,
GerardM
It is certainly true that many subjects that "merit" an article do not have an article. There may be an article in one language and not in another. Wikidata has the option to add a label for such an article as a place holder and while it has not been written it can show nicely red on a disambiguation page. The point here is that it is legitimate for lists and disambiguation pages to have red links. When such lists are completed in Wikidata they can easily be translated to other languages and provide basic information that may be of interest.
The red links in the list of Muhammeds are both kings of the Sayfawa dynasty. Sadly they are not even all the kings called Muhammed who are part of the Sayfawa dynasty with a red link. Just consider what would happen when disambiguation lists are presented from Wikidata; it makes it easier to start articles because relevant information may be available thanks to work done in another language.
When you consider the options, all "red links" could be known to Wikidata. As a result you can complete all lists without having to write articles and, you will deal with disambiguation issues sooner rather than later and, why have wikilinks when integrity is better maintained in Wikidata anyway?
Thanks,
GerardM
Sunday, April 28, 2013
Spouse does not translate
I love #Wikidata, it gets so many things right. At the same time it fails by design. Consider; there is this article about a woman. In English she may have a spouse. In translation there is a problem; because the spouse has a spouse and it is a different word depending on the gender of the person involved.
This is the kind of problem that has been solved for the MediaWiki software. It means that you consider the sex of the person involved. Applying this principle on Wikidata is possible because we do register if a person is male or female.
It seems obvious; sex matters. In many languages you address people based on their sex. As long as Wikidata is not able to address this issue, it is broken. When this is intentional, it is broken by design.
Thanks,
GerardM
This is the kind of problem that has been solved for the MediaWiki software. It means that you consider the sex of the person involved. Applying this principle on Wikidata is possible because we do register if a person is male or female.
It seems obvious; sex matters. In many languages you address people based on their sex. As long as Wikidata is not able to address this issue, it is broken. When this is intentional, it is broken by design.
Thanks,
GerardM
Thursday, April 25, 2013
#OmegaWiki works and so can #Wikidata
OmegaWiki provides functionality that is on the agenda for Wikidata. The OmegaWiki community has ALWAYS wanted to be a Wikimedia project.
What it already provides is:
What it already provides is:
- links to Wikipedia articles in other languages when the article does NOT exist in the preferred language
- links to a commons category associated with a concept
- a picture painting the thousand words associated with the concept
Yes, it is quite shocking; it also provides multi lingual dictionary support. OmegaWiki has even been described in a book on the development of lexicography.
Both OmegaWiki and Wikidata are "right in front of us" and, the functionality described above is imho "right for us". The challenge we face is to do something that is "shockingly rare"; meet halfway.
The road towards such a meeting could be:
- Adopt OmegaWiki by the WMF in a labs kinda construction
- Create the logical requirements of OmegaWiki data in Wikidata
- Convert the OmegaWiki content to Wikidata technology
- See how it fits in the big picture
- Make it fit in the big picture
We do not need much talk and, the amount of work needed is relatively minor. The benefits are legion because even when all the OmegaWiki data is not used in the end, the lessons learned will be.
Thanks,
GerardM
Machine translation for #Wikipedia
The Wikimedia Foundation suggests that machine translation is the kind of infrastructure that makes sense to it. Given what it aims to do: making knowledge available to everyone, this makes perfect sense. A lot of translation has already been going on in order to fill many gaps in the many Wikipedias and machine translations were often an important part of this.
One of the arguments why the WMF could enter the fray is that it has something to add. It does have monetary reserves but more importantly it has several resources that may make a difference. The biggest two are Wikipedia itself and the other are its awesome communities.
When translating a Wikipedia article, the concepts that are specific to a subject are likely to be found in that article. Similarly when such concepts have their own article, they will contain a similar set of concepts. Combine this with a multi-lingual dictionary build with Wikidata technology along OmegaWiki lines and it will be relatively easy to find the corresponding expressions in articles in different languages on the same subject.
The point here is that the meanings of words do not exist in a vacuum.
When such concepts have been identified and linked to Wikipedia articles and dictionary meanings it becomes possible to help people understand a text in a different language by providing native language support.
<GRIN> I know Erik Moeller has a lot of experience in this field and I know mutual friends are quite interested to help </GRIN>
Relevant is that we do not have to invent something new; it has been part and parcel of things we have done before. The difference is that we gained in experience and, technology has evolved as well.
Thanks,
GerardM
One of the arguments why the WMF could enter the fray is that it has something to add. It does have monetary reserves but more importantly it has several resources that may make a difference. The biggest two are Wikipedia itself and the other are its awesome communities.
When translating a Wikipedia article, the concepts that are specific to a subject are likely to be found in that article. Similarly when such concepts have their own article, they will contain a similar set of concepts. Combine this with a multi-lingual dictionary build with Wikidata technology along OmegaWiki lines and it will be relatively easy to find the corresponding expressions in articles in different languages on the same subject.
The point here is that the meanings of words do not exist in a vacuum.
When such concepts have been identified and linked to Wikipedia articles and dictionary meanings it becomes possible to help people understand a text in a different language by providing native language support.
<GRIN> I know Erik Moeller has a lot of experience in this field and I know mutual friends are quite interested to help </GRIN>
Relevant is that we do not have to invent something new; it has been part and parcel of things we have done before. The difference is that we gained in experience and, technology has evolved as well.
Thanks,
GerardM
Sunday, April 14, 2013
#Wikivoyage #statistics revisited
I am happy to report that many of the issues with several statistics have been resolved. The most visible improvement is that the page views for Wikivoyage are now available and updated on a daily basis. When you check out these statistics you notice that Wikivoyage got the most attention when it was started. Now in the fourth month we can notice some stabilisation.
The latest Wikivoyage projects in Hebrew and Ukrainian are now included as well and it will be interesting to follow how they will do in the future.
I am really happy that many issues have been resolved. If anything, statistics are trusted because of the consistency and quality of the presentation. When you follow developments regularly you get a feeling for the underlying data. Waiting for the resolution of what seems like cosmetic issues destroys that feeling.
Thanks,
GerardM
Wednesday, April 10, 2013
Sign languages are important everywhere
When in 2005 the Austrian Sign language was constitutionally acknowledged, there was a curious second sentence added: "The Austrian Sign language is acknowledged. All details are determined by law". This sentence was a puzzle from the begin, because either it was trivial (then now one would add it because it should be found after every single point of the constitution) or there was some hidden intent behind it.
Now there is proof for the hidden intent by letters coming from "horrible jurists" of two ministries: They tell us that since there is no respective law, Austrian Sign Language cannot be established as a mother tongue in schools for deaf pupils. The jurists use the tardiness of the education ministry (this would have been obliged to develop realization laws for deaf people) as an argument that the Austrian government now dismisses the language rights of deaf people in the education process.
I read this news and, I am sure that these jurists have no clue why it is so important for children to learn to read and write in their mother tongue.. It is beneficial for all of their academic career. This is known to be true in the USA, Tunisia and Saudi Arabia ... it is also true for Austrians.
Thanks,
Gerard
Now there is proof for the hidden intent by letters coming from "horrible jurists" of two ministries: They tell us that since there is no respective law, Austrian Sign Language cannot be established as a mother tongue in schools for deaf pupils. The jurists use the tardiness of the education ministry (this would have been obliged to develop realization laws for deaf people) as an argument that the Austrian government now dismisses the language rights of deaf people in the education process.
I read this news and, I am sure that these jurists have no clue why it is so important for children to learn to read and write in their mother tongue.. It is beneficial for all of their academic career. This is known to be true in the USA, Tunisia and Saudi Arabia ... it is also true for Austrians.
Thanks,
Gerard
Thursday, March 28, 2013
I want #Wikimedia #statistics I can rely on
Erik Moeller asked the question; "what statistics do you like best, do you think most relevant". My answer is: statistics I can rely on.
There are several developers working on statistics. Many more than several years ago. I am sure that there is a need for data driven development and I am sure that the underlying numbers for many projects require time and effort. It does not mean that the statistics that are updated on a daily basis can be neglected.
Today I learned that two new Wikivoyage projects were created; one in Hebrew and one in Ukrainian. And after more than a month I checked the statistics for pageviews for Wikivoyage. They are still very much broken. There are other bugs that I reported as well. I cannot be bothered to check their continued existence.
Erik really, having statistics is wonderful. Having many statistics is evn better. It does not matter at all if you cannot rely on them. That way they become the infamous mantra of "Lies, damned lies and statistics".
Hmmm, I am complaining here. Check out my blog; I use statistics frequently to report on things Wiki.
Thanks,
GerardM
#ASL; the most challenging language at #Translatewiki II
An effort is underway to localise American Sign Language at translatewiki.net. It is a challenge and at this stage, it works by copy and pasting translations from Signpuddle. The result at twn looks really weird; it is a string of numbers.
To learn if these numbers actually work, I copied the localisation for Wednesday to my userpage on the ase.testproject. It did not work. The copied content did not create something in SignWriting. What I got were some numbers and to make it work I needed something else..
Not only knowing that something is in SignWriting but also what sign language it is will be really powerful.
Thanks,
GerardM
To learn if these numbers actually work, I copied the localisation for Wednesday to my userpage on the ase.testproject. It did not work. The copied content did not create something in SignWriting. What I got were some numbers and to make it work I needed something else..
- M512x531S2e508489x504S18600492x470
- <signtext clear=0>M512x531S2e508489x504S18600492x470</singtext>
- <lang=ase-Sgnw>M512x531S2e508489x504S18600492x470</lang>
Not only knowing that something is in SignWriting but also what sign language it is will be really powerful.
Thanks,
GerardM
Saturday, March 23, 2013
#ASL; the most challenging language at #Translatewiki
A #Wikipedia in "our" language is what people are working towards in the incubator. There are many challenges that need to be taken before this dream is fulfilled. Messages need to be localised, articles need to be written and finally the text has to be proven to be in "my" language. Several challenges that take time and effort, all for the big moment when one more Wikipedia is created.
They are challenges but they are dwarfed by the challenges facing the people working towards a Wikipedia in American Sign Language. First of all, their effort cannot be part of the incubator. Their incubator project is in a Wikimedia Labs environment. An environment with experimental software that allows them to write their own language.
The SignWriting script is written from top to bottom as you can see in the screenshot. This is not supported by MediaWiki. The SignWriting script is not yet part of Unicode. MediaWiki only supports scripts that are encoded in Unicode.
Consider what it does to localising the software at translatewiki.net. Below you find what Wednesday looks like for Adam, who is working on the localisation of the most used messages. It must be copying data from one environment to the next, not even knowing what the effect will be when it finally shows in the labs environment.
If anything I know how much hard work goes into new projects. The effort for the Wikipedia in American Sign Language is probably the most complicated of them all.
Thanks,
GerardM
They are challenges but they are dwarfed by the challenges facing the people working towards a Wikipedia in American Sign Language. First of all, their effort cannot be part of the incubator. Their incubator project is in a Wikimedia Labs environment. An environment with experimental software that allows them to write their own language.
The SignWriting script is written from top to bottom as you can see in the screenshot. This is not supported by MediaWiki. The SignWriting script is not yet part of Unicode. MediaWiki only supports scripts that are encoded in Unicode.
Consider what it does to localising the software at translatewiki.net. Below you find what Wednesday looks like for Adam, who is working on the localisation of the most used messages. It must be copying data from one environment to the next, not even knowing what the effect will be when it finally shows in the labs environment.
If anything I know how much hard work goes into new projects. The effort for the Wikipedia in American Sign Language is probably the most complicated of them all.
Thanks,
GerardM
Friday, March 22, 2013
#FUEL makes for standards in localisations
#RedHat and #Wikimedia Foundation engineers have been collaborating on providing support for the languages of India for some time now. However, both Red Hat and the Wikimedia Foundation support people from all over the world. Both organisations know the perils of internationalisation intimately well.
As this cooperation is progressing nicely, one of the early results can be found at translatewiki.net. This is where MediaWiki is localised and this is where software is localised in more languages than anywhere else. What you will find is that the Fuelproject can now be localised in other languages than the languages from India.
This project "aims at solving the Problem of Inconsistency and Lack of standardization in Software Translation across the platform". For this to work optimally, such an approach should not be restricted to India but it should be aimed to any and all languages.
How the terminology defined in FUEL will be used in translatewiki.net is not yet clear. What is obvious is the intention to standardise terminology as much as possible. This will make using software more predictable and that is very much to the benefit of everyone.
Thanks,
GerardM
As this cooperation is progressing nicely, one of the early results can be found at translatewiki.net. This is where MediaWiki is localised and this is where software is localised in more languages than anywhere else. What you will find is that the Fuelproject can now be localised in other languages than the languages from India.
This project "aims at solving the Problem of Inconsistency and Lack of standardization in Software Translation across the platform". For this to work optimally, such an approach should not be restricted to India but it should be aimed to any and all languages.
How the terminology defined in FUEL will be used in translatewiki.net is not yet clear. What is obvious is the intention to standardise terminology as much as possible. This will make using software more predictable and that is very much to the benefit of everyone.
Thanks,
GerardM
#MediaWiki powers #jQuery for language support
#Wikipedia is powered by MediaWiki. Wikipedia has a version in over 280 languages and, more languages are waiting in the wings. All these languages need support. This support is needed in all the computer languages used for Wikipedia; in PHP, Javascript even Lua.
With so many languages to support, you really need a standardised way to provide the support. Support for localisation, for input methods, for fonts.
One of the really brilliant software projects provides language support in jQuery and consequently in Javascript. It is under active development and consequently more and more languages are supported in this way. Currently there are over 155 input methods for over 75 languages - it is already the largest repository of input methods on the Web and, it is open source ...
This development has not gone unnoticed. People who support languages and scripts not supported by the Wikimedia Foundation have a hard time finding support for their languages, for their fonts and input methods. I have seen an implementation for the Thuɔŋjäŋ languages developed by Andrew Cunningham.
Even better, Andrew was happy enough to talk to Martin and this may result in even more collaboration on languages that deserve the same support just like any other language.
Do you want to experience the full power of jQuery.ime like me? We just have to wait for the Wikimedia Foundation to roll it out to its own websites. It is great dogfood, they just have to eat it.
Thanks,
GerardM
With so many languages to support, you really need a standardised way to provide the support. Support for localisation, for input methods, for fonts.
One of the really brilliant software projects provides language support in jQuery and consequently in Javascript. It is under active development and consequently more and more languages are supported in this way. Currently there are over 155 input methods for over 75 languages - it is already the largest repository of input methods on the Web and, it is open source ...
This development has not gone unnoticed. People who support languages and scripts not supported by the Wikimedia Foundation have a hard time finding support for their languages, for their fonts and input methods. I have seen an implementation for the Thuɔŋjäŋ languages developed by Andrew Cunningham.
Even better, Andrew was happy enough to talk to Martin and this may result in even more collaboration on languages that deserve the same support just like any other language.
Do you want to experience the full power of jQuery.ime like me? We just have to wait for the Wikimedia Foundation to roll it out to its own websites. It is great dogfood, they just have to eat it.
Thanks,
GerardM
Monday, March 18, 2013
Endowment and #Wikimedia
Every so often, there is talk about an endowment fund for the Wikimedia Foundation. This means that a huge pile of money is kept is reserve for a rainy day.
There are a few problems as far as I am concerned. It makes sense when there is a reasonable fear that we will not be able to fund the projects of the WMF in the future. It assumes all the things the WMF could fund are being funded and are competently managed. The return on investment of putting money in an investment fund .. eh .. endowment fund are better than the return on investment of putting this money to work. People who administer investment funds are bankers.
Given the success of the fundraisers of the WMF, the question if the WMF cannot raise enough funds for its current activities is a joke.
What the WMF currently does and what it could do are two different things. Answering the second can only be answered by looking at what the WMF is there for. I refer to the mission statement on Meta.
When I look at this text, the last two words can be seen a reason for an endowment.
However, I am certain that we can do more to "empower and engage people around the world". Just consider how much we do in the United States and compare it to Africa, Asia and South America combined. The WMF works in collaboration with a network of chapters.. Most countries do not have a chapter or another organisation .. yet.
Given that realistically the WMF is only investing in Wikipedia, it can not be said that we are doing everything possible to realise the mission statement. It can be argued that we do everything we can manage at this time.
There is enough room for growth and much of it does not need to cost us anything but time and effort. What is needed is the coordination, the planning and the will to go where Wikimedians have not gone before.
Thanks,
GerardM
PS for an endowment to work well you have to trust and pay a banker. Personally I have more trust in us spending money wisely.
There are a few problems as far as I am concerned. It makes sense when there is a reasonable fear that we will not be able to fund the projects of the WMF in the future. It assumes all the things the WMF could fund are being funded and are competently managed. The return on investment of putting money in an investment fund .. eh .. endowment fund are better than the return on investment of putting this money to work. People who administer investment funds are bankers.
Given the success of the fundraisers of the WMF, the question if the WMF cannot raise enough funds for its current activities is a joke.
What the WMF currently does and what it could do are two different things. Answering the second can only be answered by looking at what the WMF is there for. I refer to the mission statement on Meta.
The mission of the Wikimedia Foundation is to empower and engage
people around the world to collect and develop educational content under
a free license or in the public domain, and to disseminate it effectively and globally.
In collaboration with a network of chapters, the Foundation provides the essential infrastructure and an organizational framework for the support and development of multilingual wiki projects and other endeavors which serve this mission. The Foundation will make and keep useful information from its projects available on the Internet free of charge, in perpetuity.
In collaboration with a network of chapters, the Foundation provides the essential infrastructure and an organizational framework for the support and development of multilingual wiki projects and other endeavors which serve this mission. The Foundation will make and keep useful information from its projects available on the Internet free of charge, in perpetuity.
When I look at this text, the last two words can be seen a reason for an endowment.
However, I am certain that we can do more to "empower and engage people around the world". Just consider how much we do in the United States and compare it to Africa, Asia and South America combined. The WMF works in collaboration with a network of chapters.. Most countries do not have a chapter or another organisation .. yet.
Given that realistically the WMF is only investing in Wikipedia, it can not be said that we are doing everything possible to realise the mission statement. It can be argued that we do everything we can manage at this time.
There is enough room for growth and much of it does not need to cost us anything but time and effort. What is needed is the coordination, the planning and the will to go where Wikimedians have not gone before.
Thanks,
GerardM
PS for an endowment to work well you have to trust and pay a banker. Personally I have more trust in us spending money wisely.
Help #Wikipedia Zero and localise the #mobile software
Learning about the current state of Wikipedia Zero can be done through this presentation. It is well worth it. The things that I found most interesting are:
- The use of mobile phones is different from what was expected
- The Wikipedia Zero project has a measurable and positive impact
- There are things that can be improved upon, things like localisation
What many people do not realise is that the localisation of the mobile software is not part of the MediaWiki software. Traffic from mobile devices is what grows traffic for Wikipedia a lot. The experience is that when the software has been properly localised, it has a big impact.
In India we have had multiple localisation drives. Every time we ask the same things;
- please do the most used messages first, they are what people see the most
- then do the rest of the MediaWiki messages
- The extensions used by Wikipedia are the next lot that need doing
- Yes, there are all these other MediaWiki messages as well
Localisation has a clear impact on the number of readers and that is what ultimately grows the community.
So what we really, really need are for the mobile messages to be localised as well because this is where the new readers for our projects can be found.
Thanks,
GerardM
Friday, March 15, 2013
#Pope Francis on #Wikidata
The good news of Wikidata is that it may become the repository of data on many subjects. The bad news; it can get it wrong.
Pope Francis is the new pope of the Roman Catholic church. He is not the pope of Catholicism. This is rather basic.
There are all kinds of identifiers associated. Ever heard of a "Viaf identifier", a "GND identifier", "SUDOC" or a "GND identity type" and really should they be so prominently displayed?
The picture is not of the pope, it is of the former cardinal. The coat of arms is the coat of arms of the former cardinal..
The picture in my blog is of pope Francis; he is smiling.. And yes, I can edit Wikidata.
Thanks,
GerardM
Pope Francis is the new pope of the Roman Catholic church. He is not the pope of Catholicism. This is rather basic.
There are all kinds of identifiers associated. Ever heard of a "Viaf identifier", a "GND identifier", "SUDOC" or a "GND identity type" and really should they be so prominently displayed?
The picture is not of the pope, it is of the former cardinal. The coat of arms is the coat of arms of the former cardinal..
The picture in my blog is of pope Francis; he is smiling.. And yes, I can edit Wikidata.
Thanks,
GerardM
Saturday, March 02, 2013
Using #Wikidata in #OmegaWiki
Wikidata knows about the existence of #Wikipedia articles. One really nice side effect is that OmegaWiki will point you to the article in your language, Arabic in the example, even when OmegaWiki does not have a translation in your language.
Thanks,
GerardM
Thursday, February 28, 2013
A challenge for #Wikidata
Not only #Wikipedia but also #Wikisource and #Wikiquote are in need for the #Wikidata treatment. OmegaWiki can demonstrate why. Julius Caesar has an article in all the languages with a translation. All that is needed is a link to Wikidata where Julius is known as "Q1048".
Linking to the quotes of Caesar on Wikiquote is not so easy; it has to be done for each language to the Wikiquote where there are quotes by Caesar. This is also needed for Wikisource for books or other publications. You may notice that Julius Caesar is a writer and a person. A writer publishes and any person may produce quotable quotes.
The challenge for Wikidata is to go the extra mile and rid us from those pesky interwiki links... on all the Wikimedia projects. These links can in return be used on all the projects. It also makes searching more interesting... Julius Caesar; what do you want? An article about him, his sources, his quotes..
What we at OmegaWiki would like is for Wikidata to go even further and incorporate our data. This will be a big benefit because it will allow people to search in Nepali for जुलियस सिजर and have a similar result at no extra cost as in English.
Thanks,
GerardM
Denny, I do not agree
#Wikidata is a wiki and consequently Caesar can have a capitol according to Denny's blogpost. While I applaud the sentiment, it is not really practical.
Denny is the big man for Wikidata and he is a human and, he can be both right and wrong. So let me suggest why I do not support him in this and, let me suggest a compromise.
The one thing that will help is when there are classes and when those classes have attributes. For instance; a country has a national anthem, a capitol, a currency, a head of state etcetera. When you state that Caesar is a country, it then follows that he has a capitol, a currency, a head of state.. For Wikipedia this is really powerful. When all the countries have been labelled as such, it is easy and obvious to fill in the missing values for use in an info-box. The most valuable part of it is, that these values can be translated for use in info-boxes in other languages. Yes, the values can have their own Wikipedia articles as well as a label to use in an info-box.
When a label is used in an info-box, in any language, these labels are obviously of higher value than the labels that are not used. So Denny can have his way and as long as labels are values when they are used in a Wikipedia all the extra labels will not be very much in the way. Hard-drives are cheap. It is possible to have classes with associated labels. These classes can be a associated with a particular info-box.
This is what Julius Caesar looks like at OmegaWiki. As you can see it is linked to Wikidata, it is linked to Wikipedia and, it is linked to Commons. It would be so cool when we have the best of both worlds in a joined project.
Please Denny, lets cross the Rubicon.
Thanks,
GerardM
Denny is the big man for Wikidata and he is a human and, he can be both right and wrong. So let me suggest why I do not support him in this and, let me suggest a compromise.
The one thing that will help is when there are classes and when those classes have attributes. For instance; a country has a national anthem, a capitol, a currency, a head of state etcetera. When you state that Caesar is a country, it then follows that he has a capitol, a currency, a head of state.. For Wikipedia this is really powerful. When all the countries have been labelled as such, it is easy and obvious to fill in the missing values for use in an info-box. The most valuable part of it is, that these values can be translated for use in info-boxes in other languages. Yes, the values can have their own Wikipedia articles as well as a label to use in an info-box.
When a label is used in an info-box, in any language, these labels are obviously of higher value than the labels that are not used. So Denny can have his way and as long as labels are values when they are used in a Wikipedia all the extra labels will not be very much in the way. Hard-drives are cheap. It is possible to have classes with associated labels. These classes can be a associated with a particular info-box.
This is what Julius Caesar looks like at OmegaWiki. As you can see it is linked to Wikidata, it is linked to Wikipedia and, it is linked to Commons. It would be so cool when we have the best of both worlds in a joined project.
Please Denny, lets cross the Rubicon.
Thanks,
GerardM
Wednesday, February 27, 2013
What statistics for #Wikipedia
The question was: "What statistic inspires you most". My answer; it depends what hat I am wearing.
The statistic that means a lot to me are the page views. However, they do not truly show how many PEOPLE view the pages. There are the bots that spoil the view. But there is so much history there.. Even with the bots you get a picture of growth.
However, wearing another hat, I want to see an aggregation of views of GLAM objects for the GLAM's we cooperate with. This is so vital for making the point that Wikipedia is making cultural heritage visible. It is the one argument everyone "understands".
Yes, there are many languages but for this I want to know how well does MediaWiki work for a language and this can be found best in the stats at translatewiki.net.
Really Erik, there is no one statistic that is best. There are horses for courses. What is vital that the statistics can be trusted. If anything, improved reliability of the data is what would help us most.
Thanks,
GerardM
The statistic that means a lot to me are the page views. However, they do not truly show how many PEOPLE view the pages. There are the bots that spoil the view. But there is so much history there.. Even with the bots you get a picture of growth.
However, wearing another hat, I want to see an aggregation of views of GLAM objects for the GLAM's we cooperate with. This is so vital for making the point that Wikipedia is making cultural heritage visible. It is the one argument everyone "understands".
Yes, there are many languages but for this I want to know how well does MediaWiki work for a language and this can be found best in the stats at translatewiki.net.
Really Erik, there is no one statistic that is best. There are horses for courses. What is vital that the statistics can be trusted. If anything, improved reliability of the data is what would help us most.
Thanks,
GerardM
Wednesday, February 06, 2013
More traffic for the smaller #Wikipedia projects
The biggest Wikipedia in pageviews is the English language Wikipedia. Last month it generated 49.36% of all the traffic. This left 40.34% for the numbers 2 to nine and 10.3% for all the other Wikipedias.
Traffic is growing really well; 19% on a year to year basis for all the Wikipedias and 20% for the English Wikipedia. What the chart to the left shows is that slowly but surely the "other" languages are doing better than the top 10.
The current picture will change as it did in the past. Consider that the Dutch Wikipedia was bigger than the Chinese Wikipedia. The Dutch is growing at 16% while the Chinese is powering up the ladder with a healthy 46%.
There are many languages that have the potential to move on in the traffic rankings. This may be because the content in those languages improves or because the infrastructure in a country changes. I have reliable information that in India mobile traffic went from 55m to 85m in a matter of 6 months and this while it is too soon to attribute this to Wikipedia Zero. This is not obvious from the statistics published for India.
In a way I have been cheating; the numbers in the charts above exclude mobile traffic. This is because these numbers make it more obvious how well the "other" languages are doing. However, as you can see something similar is happening when mobile traffic is included.
With Wikipedia Zero growing in relevance, it seems obvious that content that is relevant to the people who use this service will grow in demand. As this is addressed in the languages people read, it will mean that the growth of popularity of Wikipedia articles will continue for some time. I also expect that the "other" languages will become more relevant and assertive.
Thanks,
GerardM
Traffic is growing really well; 19% on a year to year basis for all the Wikipedias and 20% for the English Wikipedia. What the chart to the left shows is that slowly but surely the "other" languages are doing better than the top 10.
The current picture will change as it did in the past. Consider that the Dutch Wikipedia was bigger than the Chinese Wikipedia. The Dutch is growing at 16% while the Chinese is powering up the ladder with a healthy 46%.
There are many languages that have the potential to move on in the traffic rankings. This may be because the content in those languages improves or because the infrastructure in a country changes. I have reliable information that in India mobile traffic went from 55m to 85m in a matter of 6 months and this while it is too soon to attribute this to Wikipedia Zero. This is not obvious from the statistics published for India.
In a way I have been cheating; the numbers in the charts above exclude mobile traffic. This is because these numbers make it more obvious how well the "other" languages are doing. However, as you can see something similar is happening when mobile traffic is included.
With Wikipedia Zero growing in relevance, it seems obvious that content that is relevant to the people who use this service will grow in demand. As this is addressed in the languages people read, it will mean that the growth of popularity of Wikipedia articles will continue for some time. I also expect that the "other" languages will become more relevant and assertive.
Thanks,
GerardM
Sunday, February 03, 2013
#Wikivoyage #statistics II
The first month of Wikivoyage statistics is in. The page views statistics are now regularly updated. There is now a baseline to compare future months with. Last week I blogged about Wikivoyage statistics and at that time traffic was expected to reach 12.8 M. As you can see in the screen dump, it ended up at 17.3 M.
Over time Wikivoyage will become increasingly Wikivoyage. It will slowly but surely morph into something that can be distinguished from Wikitravel. Once that process is well under way, I am sure that there will be a growing public for this Wikimedia project.
Thanks,
GerardM
Saturday, February 02, 2013
#CLDR gets the sorting right
When you sort, order will be created in the predetermined way. Another word for such a predetermined way is called the "collation order". When you sort tea, you make sure that only the tea leaves of sufficient quality are left. The characters in the words determine where the word can be found in a sorted list.
The collation order is a standard and, the CLDR is the name of the standard. Unicode, the organisation behind the CLDR has made a big change in the order. From now on, the character of the script of the language take precedence over the Latin script.
This change affects many languages and, there is a document mentioning them all.
Thanks,
GerardM
Monday, January 28, 2013
#Wikivoyage #statistics
Page views are now also collected for Wikivoyage. At this stage it is indicated that they are page views from non-mobile devices. It is likely that the mobile page views will follow at a later date; Wikivoyage is after all the resource for people who are travelling or are about to travel.
Thanks,
GerardM
Wednesday, January 23, 2013
Give #Europeana more of a global appeal
The Europeana blog enthuses about "Europioneers" and asks if you know anyone who you would like to nominate someone as one of "Europe’s finest technology entrepreneurs."
Europeana is a great initiative. It brings European culture to the Internet and it is becoming increasingly relevant and useful. Culture is an export product and at one time everything was ready to make the Europeana software available in Chinese and other languages that are not official languages of the EU.
This is still a great idea and the software had been localised in many languages. Several languages were fully localised and then, then it was decided by someone "higher up" that it was not such a good idea after all.
Pioneering is about breaking ground and, translatewiki.net did this for the world, Europe and Europeana. Given that Europe's finest technology entrepreneurs can be nominated, I nominate Siebrand Mazeland not only for his unrecognised work for Europeana but also for the important work he and his team do for the internationalisation and localisation of software and for the platform where people actually can do this work.
Thanks,
GerardM
Tuesday, January 22, 2013
#Wikipedia traffic includes #mobile
Once upon a time, there was the encyclopaedia that could. It lived on the Internet and people approached it from their laptop or computer. Statistics showed how it grew in the different languages. At a later date there was this application in Ruby on Rails that provided a better experience for people reading Wikipedia on mobile phones. It was significantly different and it showed great potential. The statistics for mobiles were available separately. There was even a statistics page combining the two.
Nowadays, the Wikipedia experience on mobile phones has improved a lot. As a result the mobile traffic in January 2012 is 14.7% of total traffic and, the traffic from mobile phones is responsible for most of the growth in traffic for Wikipedia.
It is quite reasonable to show the all inclusive statistics from the Wikipedia statistics page. Obviously the breakdown in traffic from mobile and non-mobile is relevant. They are important for the people who care for these numbers. The Wikipedia traffic does include traffic from mobile and the combined statistics show most clearly that the appetite for quality information is still growing after all these years.
Thanks,
GerardM
Friday, January 18, 2013
#Depression and Aaron and all the others
A lot has been said about the untimely death of Aaron Swartz. It is sad that a life with so much promise comes to an early end. I met Aaron at Wikimania, he presented there as well.
Aaron like me and many other Wikimedians suffered from depression. The circumstances were not favourable for him; there was a real chance for him to go to jail. He could not cope with that possibility, he committed suicide.
Depression is a killer. For the people prosecuting Aaron, his case was business as usual. Now that he is dead, they have come to realise that Aaron was anything but usual.
Many of the people who read and edit Wikipedia suffer from all kinds of psychiatric disorders. For many of them being a Wikimedian is therapeutic.
If anything I wish that more people realise that many people who suffer from psychiatric disorders do contribute to the common good.
Thanks,
Gerard
Aaron like me and many other Wikimedians suffered from depression. The circumstances were not favourable for him; there was a real chance for him to go to jail. He could not cope with that possibility, he committed suicide.
Depression is a killer. For the people prosecuting Aaron, his case was business as usual. Now that he is dead, they have come to realise that Aaron was anything but usual.
Many of the people who read and edit Wikipedia suffer from all kinds of psychiatric disorders. For many of them being a Wikimedian is therapeutic.
If anything I wish that more people realise that many people who suffer from psychiatric disorders do contribute to the common good.
Thanks,
Gerard
Thursday, January 17, 2013
#Wikidata is used on the Hungarian #Wikipedia
The Wikipedia in Hungarian is the first Wikipedia that makes use of Wikidata for its interwiki links. This is great news because the bots that update these links are no longer needed. This will make it much easier to resolve the many issues that exist with the old interwiki approach.
The way to update interwiki links on the Hungarian Wikipedia is now by updating Wikidata. Recently many of the messages that support Wikidata have been localised in Hungarian making it much easier.
I hope that many Hungarians will find their way to Wikidata and make it a success. It will be awesome once the old interwiki links have gone the way of the dodo. In preparation for that day, you can check out if there are messages in your language that need localisation.
Thanks,
GerardM
Wednesday, January 16, 2013
Great news from the #ngadc, they dig the public domain!
" With the launch of NGA Images, the National Gallery of Art implements an open access policy for digital images of works of art that the Gallery believes to be in the public domain. Images of these works are now available free of charge for any use, commercial or non-commercial. Users do not need to contact the Gallery for authorization to use these images. They are available for download at the NGA Images website (images.nga.gov). See Policy Details below for specific instructions and notes for users."This is the kind of news that makes for happy reading. It provides a basis for happy collaboration that starts with the GLAM people in our Wikimedia movement. When you read the terms of use for the images that are in the public domain, they make sense.
The realisation that public domain is best expressed by providing the public with access to this important resource. Once the GLAM people have a working relation, there are so many more things that can be done. This is a great step forward and the world is a better place because of it.
Thanks,
GerardM
Monday, January 14, 2013
Illustrated #Wikispecies Tree
Wikispecies is one #Wikimedia project that is probably the least well known. It exists only in one language; English. It aims to include all the species and all the other levels of the nomenclature of plants and animals.
The nomenclature is a system that relates many different levels together. As the information is available in Wikispecies and, it is possible to present this in a different way. The resulting "tree of life" is really interesting and it deserves attention. Read the announcement on the Wikispecies village pump below and check out the illustration of all apes at the bottom.
Thanks,
GerardM
The nomenclature is a system that relates many different levels together. As the information is available in Wikispecies and, it is possible to present this in a different way. The resulting "tree of life" is really interesting and it deserves attention. Read the announcement on the Wikispecies village pump below and check out the illustration of all apes at the bottom.
Thanks,
GerardM
I've created an interactive illustrated tree of life based on the Wikispecies data. You can check it out here. (See Village_Pump#Taxon_Tree, especially EncycloPetey's comments, for discussion of what such a tree can and cannot be.)
Apart from being fun to play with, one purpose this can serve is to help spot erroneous links on Wikispecies. For example, the plant family Melastomataceae is currently listed as containing Dionycha, which is a group of spiders. I've already corrected a bunch of these errors locally, so they're no longer visible in the tree linked above. However I wasn't sure how to go about correcting them on Wikispecies. Presumably what the person who entered "Dionycha" intended was something like "Dionycha_(Plantae)", so I shouldn't just delete the link, right? In any case, here is a list of erroneous links I noticed.
If you're interested in using this as a tool for verification, here is an uncorrected version based on the most recent dump (2013.01.05). Some further issues are also discussed there.
-- Lifetree (talk) 03:41, 9 January 2013 (UTC)
The #Wikimedia #Commons featured pictures
I have counted to ten .... waited some more, and more ... and am not happy with the Featured Pictures practices of Commons. They are only about technical considerations and an absolutely stunning picture, a spectacular picture good enough to win the Wiki loves Monuments 2012 world wide competition fails to qualify.
On one level, I do not need to care; it is for that in-crowd to play their game. On the other hand they expect us to cherish their best in the picture of the year. In a world that is not black and white, most of their pictures are nice. Most pictures however do not impress me as much as the winning picture of Wiki loves Monuments 2012.
Multiple languages, multiple scripts in a #Wikipedia article
Many Wikipedia articles describe something "foreign". For instance Moscow is known in the Russian language as Москва. It is a best practice to mention the native name of a subject and include the pronounciation expressed in IPA as well.
When an article includes text in multiple languages, it is important to be able to identify what text is in what language. Several purposes are served in this way.
When a Wikipedia article is parsed, such tags make it possible to enable relevant technology for that language. This is of particular relevance for the "smaller" languages because a Wikipedia in that language is quite often the biggest corpus. Research of such a corpus is a lot easier when all words that are explicitly in another language are marked as such.
Adding such tags is relatively easy. To a large extend bots can be used to add them. Making use of these tags in how MediaWiki works is something else. Tagging is however the first step that needs to be taken and, there is no reason I know of why this can not be done at this time.
Thanks,
GerardM
When an article includes text in multiple languages, it is important to be able to identify what text is in what language. Several purposes are served in this way.
- provide support for language technology like web fonts and input methods
- enable data mining for the building of spell checkers.
- create frequency lists for words
When a Wikipedia article is parsed, such tags make it possible to enable relevant technology for that language. This is of particular relevance for the "smaller" languages because a Wikipedia in that language is quite often the biggest corpus. Research of such a corpus is a lot easier when all words that are explicitly in another language are marked as such.
Adding such tags is relatively easy. To a large extend bots can be used to add them. Making use of these tags in how MediaWiki works is something else. Tagging is however the first step that needs to be taken and, there is no reason I know of why this can not be done at this time.
Thanks,
GerardM
Saturday, January 05, 2013
#Wikidata needs #localisation
A lot of hard work has gone in developing Wikidata. It is getting at a stage where the localisation of the "Wikibase - client" extension (only fourteen messages) needs to be done.
Of particular interest is the localisation in Hungarian; they are likely to get the Wikidata functionality implemented as one of the first.
As always, localisation of software is done at translatewiki.net. As always there is more that you can do for your language.
Thanks,
GerardM
Of particular interest is the localisation in Hungarian; they are likely to get the Wikidata functionality implemented as one of the first.
As always, localisation of software is done at translatewiki.net. As always there is more that you can do for your language.
Thanks,
GerardM
The growth in traffic of #Wikipedia from #mobile phones
A growth of 44% compared with the previous month and a growth of 142% compared to the previous year is exceptional. The traffic of 3,709 M in December 2012 is up from 372 M in July 2010 in other words there has been a tenfold growth in less than 2 years.
It does vindicate the effort that goes in making MediaWiki look good on mobile phones.
Thanks,
GerardM
#MediaWiki for dyslexic people
#Wikipedia states: "It is believed that #dyslexia can affect between 5 and 10 percent of a given population". There are some 315,113,038 USAmericans and therefore there are something like 23,633,478 people in the USA who have a problem with Wikipedia; they find Wikipedia hard if not impossible to read.
The Website of Dyslexia USA has on its information pages a very simple drop down box. It offers 8 different background colours. Each background colour helps a different group of people with dyslexia.
It is a cheap and easy way of growing the numbers of readers of Wikipedia in the USA. Add to this the Open-Dyslexic font we are waiting for and we help them to the resource that has been there for all the rest of us.
PS It will work for people in other countries or who speak other languages too.
Thanks,
GerardM
The Website of Dyslexia USA has on its information pages a very simple drop down box. It offers 8 different background colours. Each background colour helps a different group of people with dyslexia.
It is a cheap and easy way of growing the numbers of readers of Wikipedia in the USA. Add to this the Open-Dyslexic font we are waiting for and we help them to the resource that has been there for all the rest of us.
PS It will work for people in other countries or who speak other languages too.
Thanks,
GerardM
Friday, January 04, 2013
#Dutch #government loves #WikiLovesMonuments
Obviously the RCE has its own photos of most of the monuments; they are used to do whatever is needed to maintain the necessary information. The RCE has been taking pictures for many years. Currently there are some 555.000 images in its collection. All of them are now available under the Creative Commons Attribution-Share Alike 3.0 NL license. All of them will be uploaded to Commons.
With all these images available on Commons, this collection is available to people from all over the world who want to learn about the Dutch cultural heritage and its monuments. As this is a cooperation with the RCE, it comes with references to the RCE database where much more information can be found.
Some images like the one to the right show exceptional moments like the three pictures showing the hoisting of the church tower of the church in Bodegraven in April 1973 ( 1 2 3)
Add to the over half a million pictures from the RCE the pictures of Wiki loves Monuments and there is an obvious challenge; how to find the pictures of the same object. The good news is that the RCE numbers were available on the list.. A Commons picture of the tower in Bodegraven standing tall can be found here.
Combining all the related information is relatively easy to do in a database.. Maybe, one day Wikidata will be the platform making all these connections for us.
The RCE has given us a great gift, it is for us to use it and make it available in a usable way.
Thanks,
GerardM
Subscribe to:
Posts (Atom)











































