Seeing updates go live on a Wikipedia when updates happen on Wikidata; wow! Mr Starobinski is from Switzerland so it is not really surprising when German language and French language awards are awarded.
At first a Q1730045 showed up; it is the "Karl-Jaspers-Preis" it did not have a label so I added one in French and now it shows in black. I noticed dates, so I added them all. <grin> if there is one thing the script could do is sort them by date :) This is a wonderful experience!
Thanks,
GerardM
Showing posts with label Switzerland. Show all posts
Showing posts with label Switzerland. Show all posts
Sunday, February 07, 2016
Friday, March 06, 2015
#Kiwix - getting #Labs ready for the #Wikipedia big time
Offline #Wikipedia received a big boost. It is updating monthly its images for most of the #Wikimedia projects. Most but not all. Emmanuel was asked to write up about his challenges and I am happy to share this with his permission. Developments like this make both Labs and Kiwix even more strategic to out goals.
Thanks,
GerardM
Thanks,
GerardM
Following Yuvi's and Andrew's invitation, I write this email to explain what I want to do with Labs and share with you my first experiences.
== Context ==
Most of the people still don't have a free and cheap broadband access to fully enjoy reading Wikimedia web sites. With Kiwix and openZIM, a WikimediaCH program, we have been working on solutions for almost ten years to bring Wikimedia content "offline".
We have built a multi-platform reader and have created ZIM, a file format to store web site snapshots. As a result, Kiwix is currently the most successful solution to access Wikipedia offline.
== Problem ==
However, one of the weak point of the project is that we still don't achieve to generate often enough new fresh snapshots (ZIM files). Generating ZIM snapshots periodically (we want to provide a new fresh version each month) of +800 projects needs pretty much hardware resources.
This might look like a detail but it's not. The lack of up-to-date snapshots brakes many action within our movement to advert more broadly our offer. As a consequence, too few people are aware about it reported last Wikimedia readership update. An other side effect is that every few months, volunteer developers get the idea to build a new offline reader based on the XML dumps (the only up2date snapshots we provide for now), which is near to be a dead-end approach.
== Goal ==
Our goal with Labs is to have a sustainable and efficient solution to build, one time a month, new ZIM files for all our projects (for each project, one with thumbnails and one without). This is at the same time a requirement for and a part of a broader initiative which has for purpose to increase the awareness about our "offline offer". Other tasks are for example, storing all the ZIM files on Wikimedia servers (we currently only store part of them on download.wikimedia.org) and improve their accessibility by making them more visible (WPAR has for example customised their sidebar to provide a direct access
== Needs ==
Building a ZIM file from a MediaWiki is done using a tool called mwoffliner which is a scraper based on both Parsoid & MediaWiki APIs. mwoffliner, after scraping and rewriting content, store them in a directory. At the end, the content is then self-sufficient (without online dependencies) and can be then packed in one step in a ZIM file (using a tool called zimwriterfs).
To run this software you better have:
- A little bit bandwidth
- Low network latency (lots of HTTP requests)
- Fast storage
- Pretty much storage (~100GB per million article)
- Many cores for compression (ZIM, ZIP and picture optimisation)
My guess is that we need a total of around a dozen of VMs and 1.5 TB of storage.
- Time (~400.000 articles can be dumped per day on a machine)
== Current achievements ==
We have currently 3 x-large VMs in our "MWoffliner" project:
With them we are able to provide, one time a month, ZIM for all instances of Wikivoyage, Wikinews, Wikiquote, Wikiversity, Wikibooks, Wikispecies, Wikisource, Wiktionary and a few minors Wikipedias.
Here are a few feedbacks about our first months with Labs:
- Labs is a great tool, it's fully in the Wikimedia spirit and it works.
- Support on IRC is efficient and friendly
- We faced a little bit instability in December but instances seem to be stable now
- The Documentation on wikitech wiki seems to be pretty complete, but the overall presentation is to my opinion too chaotic and stepping-in is might be easier with a more user-friendly presentation.
- Mediawiki Sementic & OpenStackManager sync/cache/cookie problems are a little bit annoying
In general, Labs does the job, we are satisfied and think this is an adapted solution to our project.
- Overall VM performance looks good although suffering from sporadic instabilities (bandwidth not available, all the processes stuck in "kernel time", slow storage).
== Next steps ==
We want to complete our effort and mirror the biggest Wikipedia projects. Unfortunately, we have reached the limits of a traditional usage of Labs. We need more quota and we need to experiment with the NFS storage because an x-large instance in not able to mirror more than 1.5 millions of articles at a time. How might that be made possible?
Thursday, November 20, 2014
#Wikimedia & Project #Gutenberg - the sum of all knowledge
"To share in the sum of all knowledge" is the vision of the Wikimedia Foundation. The Swiss chapter does understand this really well. It has adopted Kiwix, an off line reader for content that is published in the ZIM format.
Project Gutenberg is a well established organisation dedicated to the digitisation of books. Its catalogue of 50.000 public domain books is now available to everybody, everywhere and offline as well.
Thanks to a hackathon, all books are now available in the ZIM format, you can search in all the books at the same time. The best news is that not only has this work been done for a first time, it is build in such a way that it can be easily repeated.
Future deployments may include all the books of Wikisource, books from other sources and even copyrighted works as well. The point of Kiwix is that it is an enabler, it allows for the dissemination of knowledge and to achieve THAT is what our aim is.
Congratulations to the Swiss Wikimedia chapter for providing the sustained support of this valuable project.
Thanks,
GerardM
Project Gutenberg is a well established organisation dedicated to the digitisation of books. Its catalogue of 50.000 public domain books is now available to everybody, everywhere and offline as well.
Thanks to a hackathon, all books are now available in the ZIM format, you can search in all the books at the same time. The best news is that not only has this work been done for a first time, it is build in such a way that it can be easily repeated.
Future deployments may include all the books of Wikisource, books from other sources and even copyrighted works as well. The point of Kiwix is that it is an enabler, it allows for the dissemination of knowledge and to achieve THAT is what our aim is.
Congratulations to the Swiss Wikimedia chapter for providing the sustained support of this valuable project.
Thanks,
GerardM
Monday, May 12, 2014
#WMHack #Maps and #Wikidata II
This hackathon had many people with an interest in maps come together in Zurich. There were several challenges they faced; how to represent maps in a wiki, how to store them and what do we need to know about them in Wikidata. In this mix of challenges the differences between contemporary maps and historic maps feature as well.
Wikidata needs to know several specific things; it needs to know that something is a map, it needs to know the four corners of a map, the location where that map can be found and finally it is nice to know what type of map it is. More attributes are possible but this was considered the minimum for Wikidata.
The thought process about Commons was forward looking; it is going to be "Wikidatafied" and this will surely affect current practices. Information that is currently in templates will move into Wikidata and many of the galeries and categories will surely become redundant because queries will provide a more reliable and complete result.
For Wikis, current best practices were analysed and, it was found that information on a map exists in many layers. There is a base layer and on top of that you can show a contemporary or historic map. On top of it you may want to show the shapes of countries or districts. These may be sprinkled with pointers that reflect the result of a query. To finish it off, you may want to add even more that demonstrates a point made in a particular article.
All this information needs a place. It needs a special place because you may want to use a map several times. In Zurich we ended of a working example of a map that included all these complications by inserting information in a namespace. The next challenges are to make it robust and user friendly enough.
Thanks,
GerardM
Wikidata needs to know several specific things; it needs to know that something is a map, it needs to know the four corners of a map, the location where that map can be found and finally it is nice to know what type of map it is. More attributes are possible but this was considered the minimum for Wikidata.
The thought process about Commons was forward looking; it is going to be "Wikidatafied" and this will surely affect current practices. Information that is currently in templates will move into Wikidata and many of the galeries and categories will surely become redundant because queries will provide a more reliable and complete result.
For Wikis, current best practices were analysed and, it was found that information on a map exists in many layers. There is a base layer and on top of that you can show a contemporary or historic map. On top of it you may want to show the shapes of countries or districts. These may be sprinkled with pointers that reflect the result of a query. To finish it off, you may want to add even more that demonstrates a point made in a particular article.
All this information needs a place. It needs a special place because you may want to use a map several times. In Zurich we ended of a working example of a map that included all these complications by inserting information in a namespace. The next challenges are to make it robust and user friendly enough.
Thanks,
GerardM
Sunday, May 11, 2014
#WMHack #Maps and #Wikidata
This hackathon and many maps put #Zurich on the map. When you consider maps and, particularly historic maps, they have four corners and a date. That is the minimal approach to a map. You can add to this what a map intends to show, it can be a thematic map or a generic map.
When you add the four corners of a map as properties to a map, you can query for the maps that include Zurich.. When the maps are dated, you can show them in order..
It is really exciting that it has been decided what we need in a map on Wikidata.
Thanks,
GerardM
When you add the four corners of a map as properties to a map, you can query for the maps that include Zurich.. When the maps are dated, you can show them in order..
It is really exciting that it has been decided what we need in a map on Wikidata.
Thanks,
GerardM
Saturday, May 10, 2014
#WMHack - This is not a hack, we can share in the sum of available knowledge
The Wikimedia Foundation wants to share in the sum of all knowledge. What we can do now is generate text on the fly based on the knowledge available in Wikidata. In that way we can share in all available knowledge we languages on all subjects.
Wikidata may provide Wikipedia of services.When Wikidata was conceived, its first line of business was to replace all the "interlanguage links of Wikipedia. As a result it knows about more subjects than any Wikipedia. It knows about for instance more US-Americans than the English language Wikipedia. An other objective is to include statements for each item so that information can be centralised in Wikidata for use in any and all projects.
When the statements have labels in a language, it is possible to provide information in that language. It could be any language even English. The current thinking is very much: "we can serve the information boxes in articles from Wikidata". What Reasonator and WD-Search prove is that those articles do not need to exist. Most members of the South African National Assembly do not have articles in any language but information could be found in any language spoken in South Africa.
We can use machine translation to translate articles but we can also use similar algorithms to generate text based on the information we have. This has been done often in our wikiverse; they are the bot generated articles. In the Reasonator we generate text about humans in English and a few other languages. It is not rocket science to improve on what is there. In essence it is exactly what we do in our localisation functionality in translatewiki.net. It follows that we have some ability for at least 280+ languages.
This is possible with current technology, the software comes with a great pedigree. It is brought to you by the same human who started with MediaWiki. He is a scientist and the functionality is for you to enjoy in so many ways.
The point is that we can. We can share in the sum all the knowledge that is available to us. We can do more than aspire, we can share much more of the wealth that is hidden on our servers.
Thanks,
GerardM
Wikidata may provide Wikipedia of services.When Wikidata was conceived, its first line of business was to replace all the "interlanguage links of Wikipedia. As a result it knows about more subjects than any Wikipedia. It knows about for instance more US-Americans than the English language Wikipedia. An other objective is to include statements for each item so that information can be centralised in Wikidata for use in any and all projects.
When the statements have labels in a language, it is possible to provide information in that language. It could be any language even English. The current thinking is very much: "we can serve the information boxes in articles from Wikidata". What Reasonator and WD-Search prove is that those articles do not need to exist. Most members of the South African National Assembly do not have articles in any language but information could be found in any language spoken in South Africa.
We can use machine translation to translate articles but we can also use similar algorithms to generate text based on the information we have. This has been done often in our wikiverse; they are the bot generated articles. In the Reasonator we generate text about humans in English and a few other languages. It is not rocket science to improve on what is there. In essence it is exactly what we do in our localisation functionality in translatewiki.net. It follows that we have some ability for at least 280+ languages.
This is possible with current technology, the software comes with a great pedigree. It is brought to you by the same human who started with MediaWiki. He is a scientist and the functionality is for you to enjoy in so many ways.
The point is that we can. We can share in the sum all the knowledge that is available to us. We can do more than aspire, we can share much more of the wealth that is hidden on our servers.
Thanks,
GerardM
Thursday, March 27, 2014
#Monuments of #Switzerland
The problem with #data is, how do you keep it all up to date. For the monuments of Switzerland much of the data is kept on the toolserver. It works just fine.
Wikidata has matured enough to include much if not all of that same data. Missing in the official functionality are the tools to make use of it. Yes, you can store information on the connected Wikipedia articles but that is not the same as using it to administer a project like Wiki loves Monuments.
Un-official functionality meanwhile does provide much of what is needed. This is a list of all the monuments known to Wikidata in Switzerland. I added 86 pictures to the Wikidata items using the WD-Fist functionality.
As more tools get connected, these tools are increasingly attractive to use for a project. One big advantage of Wikidata is that you do not need to have an article for every monument you know about.
Thanks,
GerardM
Monday, June 04, 2012
#wmdevdays - #accessibility
At the Berlin #hackathon 2012 many people were hacking on many subjects. Kai Nissen worked on the accessibility of MediaWiki by people with a visual impairment. At the end of the hackathon several bugs were squashed and ready for review in Gerrit.
It is wonderful when reports on defects are actionable and when something gets done.
Thanks,
GerardM
It is wonderful when reports on defects are actionable and when something gets done.
Thanks,
GerardM
How was the Berlin Hackathon 2012 for you
I participated in the Hackathon for the very first time and was quite amazed about so many people coming together and actually work productively on feature enhancements or bug fixes. I had a lot of talks with people who gave really helpful feedback about the project I'm currently working on.
You have been working on accessibility for the blind ... how did you get into this subject
I was pointed to an analysis report concerning accessibility in Wikipedia that was carried out by the Swiss initiative "Access for all". While reading this I realized that a lot of the mentioned issues were
quite easy to solve. A lot of the content available on the web seems to be designed without considering accessibility aspects, although a little tweak can always have a high impact.
How do you know what to focus on
The analysis report was quite thorough and included recommendations, so it ended up to be something like a task list.
You identified a number of issues to work on this weekend .. how did it go
The issues I was working on were quite easy to fix, whenever I had problems with something there was always somebody around to help out.
Did having all these other hackers make a difference ?
The gathering of all those experienced MediaWiki developers is a really helpful thing. That applies to having certain questions answered rightaway as well as just sharing experience in whatever topic might
come up.
Brion Vibber helped you with the parser tests ...
Brion figured out what the problems was in no time. There has been a language version related bug in another patch which he simply reverted.
Are there many more accessibility issues in MediaWiki people can help with
The accessibility analysis report mentions more issues that need to be fixed to make Wikipedia and all other projects based on MediaWiki more accessible. It is clearly written and points out lacks of accessibility very
How do you continually test for good accessibility of our software
Most of the time one seldomly notices lacks of accessibility when not being affected by disabilities. Whatever issue I fixed for improving accessibility I have to keep in mind to apply that again in a similar case.
Are there best practices for coding for accessibility
There are guidelines defined by the WAI, which should be considered when coding for accessibility.
Do you have thoughts on what the Visual Editor will do for accessibility ?
Since screen readers will read what is written on the screen it also reads the wikitext as is. That might be hard to understand, especially for newcomers. Reading out a headline as a headline instead of
"equals-equals-headline- equals-equals" can be very helpful to focus on the subject itself.
I participated in the Hackathon for the very first time and was quite amazed about so many people coming together and actually work productively on feature enhancements or bug fixes. I had a lot of talks with people who gave really helpful feedback about the project I'm currently working on.
You have been working on accessibility for the blind ... how did you get into this subject
I was pointed to an analysis report concerning accessibility in Wikipedia that was carried out by the Swiss initiative "Access for all". While reading this I realized that a lot of the mentioned issues were
quite easy to solve. A lot of the content available on the web seems to be designed without considering accessibility aspects, although a little tweak can always have a high impact.
How do you know what to focus on
The analysis report was quite thorough and included recommendations, so it ended up to be something like a task list.
You identified a number of issues to work on this weekend .. how did it go
The issues I was working on were quite easy to fix, whenever I had problems with something there was always somebody around to help out.
Did having all these other hackers make a difference ?
The gathering of all those experienced MediaWiki developers is a really helpful thing. That applies to having certain questions answered rightaway as well as just sharing experience in whatever topic might
come up.
Brion Vibber helped you with the parser tests ...
Brion figured out what the problems was in no time. There has been a language version related bug in another patch which he simply reverted.
Are there many more accessibility issues in MediaWiki people can help with
The accessibility analysis report mentions more issues that need to be fixed to make Wikipedia and all other projects based on MediaWiki more accessible. It is clearly written and points out lacks of accessibility very
How do you continually test for good accessibility of our software
Most of the time one seldomly notices lacks of accessibility when not being affected by disabilities. Whatever issue I fixed for improving accessibility I have to keep in mind to apply that again in a similar case.
Are there best practices for coding for accessibility
There are guidelines defined by the WAI, which should be considered when coding for accessibility.
Do you have thoughts on what the Visual Editor will do for accessibility ?
Since screen readers will read what is written on the screen it also reads the wikitext as is. That might be hard to understand, especially for newcomers. Reading out a headline as a headline instead of
"equals-equals-headline-
--- Kai

Subscribe to:
Posts (Atom)






