Tuesday, January 29, 2008

Firefox is doing well

In a blog on Wired, the news is that Firefox is doing really well. Firefox has gained more ground on Internet Explorer. World wide 28% of the Internet users use Firefox, a new record, and 66.1% use Internet Explorer. Of interest in these numbers is that half of the IE users are still using the old IE6.

Firefox is getting ready for the release of the new Firefox 3.0. I have been using the beta 2 version for some time now and, this works out really well for me.

The best of the blog was in the end. Here they indicated that Firefox is doing so well because of the number of localisations. With 40 localisations and several languages being in beta, Mozilla is better at reaching out then Microsoft.

When this is the yardstick to measure by, MediaWiki is a league of its own. BetaWiki supports the localisation of 259 languages. More then 90% of the messages have been localised for 48 languages and for 85 languages better then 90% of the most relevant messages have been done.

What I find exciting is that particularly the languages from India are improving. A lot of work has already been done for Bengali, Telugu, Malayam and Marathi. We hope that this will remove one of the hurdles for the people from India and Bangladesh to use Commons.
Thanks,
GerardM

Monday, January 28, 2008

What if a language just does not have the word for it?

Pennsylvania Deitsch is a language spoken by some 100.000 people in North America. It does not have a separate word for Dutch, German and English. The question is how do you localise for this language when it is important to differentiate between these languages ...
Thanks,
GerardM

Sunday, January 27, 2008

The Maithili language and the Mithilakshar script

Maithili is a language spoken in India and Nepal by some 24.797.582 people. It is an official language in the Indian state of Bihar and it may be used in education.

A request was made for a Wikipedia for Maithili, it conforms to the requirements so that is not a problem. What IS a problem is that Mithilakshar, the script used to write the Maithili language, is not yet part of Unicode. The script has not even been recognised in the ISO-15924 yet.

This is the second request for a Wikipedia where the script that is used to write a language presents a problem. For modern Maithili there is the option to write in the Devangari script.
Thanks,
GerardM

Friday, January 25, 2008

Low hanging fruit

Low hanging fruit is a metaphor for easily obtainable results. When you picture this, I think of apple trees heavy with fruits and some are easy to reach. When you pick those apples, the branches become less heavy and as a result the higher up apples get out of reach. Picking the low hanging fruit means that the other results get out of reach as a consequence.

When you look at modern orchards, you see something different. The trees are no longer allowed to grow tall. All fruit have become low hanging fruit. Consequently, when you plan on some activity, and you do it in the classical way you may find low hanging fruit. When you find a way to think outside of the box, make this paradigm shift, you may have only low hanging fruit.
Thanks,
GerardM

Thursday, January 24, 2008

Maria Catlin-a

My friend Bèrto became a father of a beautiful baby girl, Maria Catlin-a. Bèrto is a happy man; his wife and his daughter are well and he does not have time to sit, drink or smoke. For a man living in Kiev this is said to be something.

Bèrto is Piedmontese; he loves his language and he wants his daughter to learn to speak Piedmontese. When you are the only one in a big place like Kiev, how do you do it. Bèrto is the guy behind the i-iter website, a website where people can call in and leave a recorded message that can be heard on the website. This service is very much to help people overcome their reluctance in voicing or writing in their language.

Bèrto told me that he is now going to ask people to record stories for Maria Catlin-a. Bedtime stories, fairy tales stories that Berto can have his daughter listen to. Stories that will help her to learn Piedmontese..

I love the idea.

Thanks,
GerardM

Wednesday, January 23, 2008

Pashto

Pashto and Dari are the two official languages of Afghanistan . Both languages are of interest from a language point of view. Dari, is considered to be the same as Farsi or Persian. Pashto is interesting for a different reason. Pashto is according to the ISO-639-3 a macro language. This means that there are three languages that all fit in under the banner of Pashto. They are Central, Northern and Southern Pashto. According to Ethnologue Southern Pashto is the Afghan National language.

In BetaWiki we try to offer the best possible user interface and consequently, when people say that they speak Pashto, we need to know what is meant by that. When it is indicated that they speak Pashto, meaning the official language of Afghanistan, we assume they write Southern Pashto. The feedback that we got that there are several dialects but that there is only one form of written Pashto.

The Ethnologue information on Northern Pashto says something different. It indicates that there is a rich literary tradition in this language. Consequently, we cannot say that written Pashto is all encompassing and this makes the use of the ps code for Pashto problematic. On the other hand the people who tell us that they speak Pashto do not appreciate that there language is split by people from the west.

What to do? The best what we can do is ask some people that are in the know. In the mean time we are happy that in BetaWiki the localisation for Pashto is improved. But we do need to know how to deal with this.
Thanks,
GerardM

Tuesday, January 22, 2008

Mingrelian

Mingrelian is a language spoken in Georgia. Ethnologue indicates that some 500.000 people speak the language. We have added this language in OmegaWiki, and the cool thing is the way the language is started. It was started with the import of some sixty language names.

As we have the language enabled on the MediaWiki level, it is now also possible to present the language names in Mingrelian. The way the MediaWiki UI is defined, is that Georgian is to be used as the fall back language. The OmegaWiki software uses the MediaWiki 1.10alpha (r26222). What we can expect when we update MediaWiki is that the localisation will be improved. Many wikis will benefit in this way.


BetaWiki is the MediaWiki wiki for localisation. It is also the place where really cool experiments are taking place. I noticed the other day that they have this cool widget where you can select your language by just starting to type its name. For me it changes the content really nicely into Dutch. Having something like this for Commons ....
Thanks,
GerardM

Sunday, January 20, 2008

Noumande (not nrm)

The Norman Wikipedia is a lively project. With 2834 articles it is doing well. Another Wikipedia, only 550 articles bigger is the Alemannic Wikipedia, another lively project. In size these projects are somewhere in the middle of the range. Both were created as a consequence of a problematic vote but it works out quite differently.

When the Alemannic Wikipedia was created, it was fully enabled in MediaWiki. This is not the case for the Norman Wikipedia, its language file only exists in the project. In order to support this project, some programming needs to be done. The BetaWiki programmers are however of the opinion, an opinion that I share, that this should not be done because the nrm code they occupy, is the code for Narom, a language spoken in Malaysia.

These two languages are treated differently. For both linguistic entities some people chose to ignore that there was no proper code. Both currently exist as a project but technically it is not possible to support them completely with the codes they abuse.

So how to deal with this. For a consistent localisation of MediaWiki it is important that we use a coding system that will not bite us. When properly implemented, the ISO-639 codes will not bite us. However it does mean that we have to fix what is broken.

On BetaWiki we have had a request for the localisation of Erzgebergish; no new project is asked for, just the localisation is requested. A request will be made, we already know that it is likely to succeed. This can also be done for Norman. When Norman is considered a dialect, it is certain that we can get a code. I do not know if Norman would be considered a language but this is something that can be discovered. With a proper code, we can have the localisation done at BetaWiki.

For Alemannic it is more complicated. It is not one language, it is several languages. What I do not know is to what extend it maps with what Ethnologue calls Alemannic ...

Thanks,
GerardM

Promoting BetaWiki

Localisation is important. It is one of the best strategies for making people feel comfortable in their Wiki project in their language. There are many projects in the Wikimedia Foundation and all deserve proper localisation. My hero of the MediaWiki localisation is Nikerabbit, who has been working on the MediaWiki localisation since 2006 first with Gangleri and now with Siebrand. He has been instrumental in starting and developing BetaWiki and it has developed in a first class environment.

The continued need for better localisation is best illustrated by numbers. According to the SiteStatistics, the WMF supports 258 languages. In the Incubator is a long list of requested projects many representing new languages waiting for a place under the sun. BetaWiki supports 257 different localisations. This is not the same as 257 languages, Chinese alone represents at least four languages.

What is also important is to consider the quality of the localisation. Only 128 of the 257 localisations have more then 50% of the most relevant messages localised and only 55 are better then 99%. For all MediaWiki messages 98 localisations do better then 50% and 56 are better then 90%. For the extensions used in the Wikimedia Foundation 31 do better then 50% and 15 do better then 90%.

When this is news to you, it may be grim reading. It reads like we are letting our readers and editors down. When you are looking for bright spots, there are plenty to find. We now have the numbers that shows our performance when it comes to localisation. We have numbers showing our progress; more languages are supported and an upwards trend can be observed in the other numbers. Personally I am happy to observe that the policy of the language committee is proving its value.

What we need is more people caring for the localisation for their language. When this localisation is done in BetaWiki, not only will people benefit from an environment that is made with them in mind, they will even get pointers indicating what messages are used for. The messages are committed daily to SVN and this makes the messages go life as soon as Brion updates the life systems.

We need to promote BetaWiki, we need to spread the message. Please help us promote BetaWiki, spread the word and check out what you can do for your language.
Thanks,
GerardM

Friday, January 18, 2008

OmegaWiki in Arabic

Yesterday I posted that the system messages of OmegaWiki were imported in Betawiki. Today I learned a new trick; how to import these messages from SVN. I am thrilled that I can show you the first results. It is wonderful that we have our content in Arabic as well.

The next thing that we really want is to show the same information as it should be shown; in a right to left orientation ... :)

Thanks,
GerardM

Wednesday, January 16, 2008

BetaWiki meets OmegaWiki

When you promote the use of BetaWiki as much as much as I do, it is surprising when the one big project I am involved in, OmegaWiki, does not use it. It is therefore that I am really happy to announce that OmegaWiki is now supported by BetaWiki for its localisation.

The messages that existed in OmegaWiki have been imported. As some OmegaWiki messages already existed in BetaWiki it is gratifying to see that some localisations already existed that we did not know about.

Thanks,
GerardM

Sunday, January 13, 2008

Localisation of the MediaWiki user interface

One of the requirements for requesting a new project within the Wikimedia Foundation is the localisation of the user interface. For a first project it was felt that we do not want to burden the new community too much and have allowed for the localisation of a subset of all the MediaWiki messages. When a subsequent project is requested for a language, we require a complete localisation of the user interface. This means all the MediaWiki messages and the messages of the extensions that are used in the Wikimedia Foundation.

It is important for the people of existing projects that the localisation is done centrally because this allows them to use their language in the user interface of Commons and other projects. Currently the localisation for many languages leaves something to be desired. This makes it harder for the people requesting new projects to fulfill this requirement.

BetaWiki has proven itself as a great facilitator for the MediaWiki localisation. It has a smooth web interface, it supports off line work by providing "gettext" or ".po file" import and export. I do urge people to support their language.

We are in the process of providing more stimuli to the localisation effort, I hope to inform you about this in the near future.

Thanks,
GerardM

Saturday, January 12, 2008

Providing information when there is little or none

The aim of Wikipedia is to provide encyclopedic information. The aim of the Wikimedia Foundation is to provide information. Consequently the goal of the WMF is broader then the goal of Wikipedia. The implications are not often considered. An other issue is that information provided is provided in a Wiki. This means that information is not necessarily complete and correct and also that information provided in one language can be and typically is substantially different in another.

When the aim is to provide information to the people of this world, it stands to reason that we want to provide the best information available. Sadly we are not able to provide the same quality information to all people in all languages because the quality will not be the same in all Wikipedias ever. When this is a given, the first line of business should be how can we provide the best available information to people. When people are looking for information in a Wikipedia, they are looking for information in their language. When there is no article, they draw a blank. This is the space where a lot of improvements are possible. This is an issue that is particularly relevant to the less resourced languages.

The first thing to appreciate is that a person often knows more then one language. This combination of languages can be anything. The challenge is to have a graceful fall back to information that is either less informative or qualitative until the point where we provide pointers to information in another source of information. The way you can move from one Wikipedia to one in another Wikipedia is by way of the "interwiki links". The information can be seen as lexical in nature with a twist. Disambiguation of homonyms is a requirement. The least information in an article is a stub. A stub can contain an "info box" and some lines of text. A stub can be written by hand in the wiki and it can be put on the Wiki by a bot. A text can be translated by hand and by machine. All these things bring there own issues. Key in the understanding is that it takes an article in order to have an interwiki link.

To reduce things to the least information we want to provide, you are left with the concept, a definition and links to information in other languages. These translations can be linked to Wikipedia articles in the languages of the translation. At this level we can provide information that is not language dependent; for instance a photo of a horse is a picture of a horse in any language. When a concept is related to other concepts, we can show these relations with the concepts preferably in the language of the reader.

Encyclopedic articles come in different states of development. A well written article on a subject in one language may be not much more then a stub in another or not exist at all. Articles can be translated by hand or by machine and in this way information can be provided. A better start for reading and editing is provided is available in this way then with a stub.

Stubs, bot created articles and translations are seen by some as problematic while others see their value. They gives rise to a constant amount of sniping with new arguments or old arguments presented as new. They distract from what we are about; we are about providing the best information we can. There is no such thing as "the" Wikipedia as there are many. Consequently the quality standards that apply to one should not be applied to another. On an intellectual level this is understood by most but regularly people find new "problems".

One recurring theme is the number of articles; people feel offended when a projects has too many bot-created articles or machine translated articles. It is felt that it is unfair; it denies the value of all the human effort that went into their projects. In many ways the arguments are similar to the ones about "Final version" and consequently the solution can be similar. When bot created articles and MT articles go into a separate namespace, they are not counted as an article. Basic information is provided and, these articles can be improved with "interwiki links" and provide a route to information in other languages. When these articles are expanded or proofread by a person, they can be moved into the main name space.

The benefit of this proposal is that we will provide more information in more languages. Most of the arguments of the exclusionists have a reasonable reply and the work people put into the creation of more information in their language has found a place.

Thanks,
GerardM

Friday, January 11, 2008

Site Matrix - BetaWiki extension of the week

BetaWiki has an extension of the week. This week it is Site Matrix. Site Matrix is a special page showing what projects exist for what languages.

It is a real eye opener. In the past many wikis have been created in anticipation of a future need, a need that never materialised. These projects have been just idling for someone to come along and start a project. In 2007 many projects that did not have any life in them were locked and deleted. There is a long list of discussions on Meta to close even more projects. I would argue that projects like the bm.wiktionary are closed as well.

I find it odd that I argue for the closure of projects...

Hmmm
Gerard

Thursday, January 10, 2008

A map with a difference

This is a map of the USA. It has been uploaded to Commons and consequently it is now freely licensed. This map has created some excitement because good educational material for the deaf is in short supply. The text that you find on the map is in ASL and the characters are SignWriting.

For American Sign Language a request has been made for a Wikipedia. The language committee of the WMF has so far not approved, nor denied it. We want to approve it; it has an ISO-639-3 code (ase), there are sufficient native speakers supporting the project... It has not been approved yet because it is technically not possible to support ASL in MediaWiki.

Putting ASL and SignWriting on the map in a Wikipedia of its own is really important. It will not only make a difference for the people who do ASL, it will also signal to people in other countries, speaking other sign languages that writing in their own language is something that can be done. This will make a real difference in the emancipation of sign languages.

It is for this reason that I am thrilled that Valerie Sutton wrote that they will be looking into the possibility of an extension to MediaWiki to make it possible and have a Wikipedia.

Thanks,
GerardM

Chemical elements - finding information on the Internet

Chemical element is a class in OmegaWiki. One hundred eighteen exist in this class and at the time of writing we have more then 12 translations for 73 languages. If every one of these languages have all translations, there would be 8614 translations, currently there are 5035 translations.

For many of these translations a Wikipedia article exists. But how can we provide information to people who do not speak one of the languages that we have a Wikipedia article for? How can we find information in the language of these people? The first port of call is Google. Google does a great job, they process daily terabytes in order to provide everyone with the ability to find what exists on the Internet. When there is information on the Internet, a word like lithium exists in at least five languages. When you are looking for information in Czech, it is likely to be drowned in information in English. A sophisticated user knows about Google's advanced search. However many languages do not exist on this list... languages like Bengali, Hindi or Swahili...

Google cannot find it yet. One of the reasons why they cannot find it is because the documents on the Internet do not tell in what language they are written. You have to have smart routines that distinguish one language from the next. It would be much easier if this information was entered with the document at the source. Sadly all web tools are not able to do this properly.

Wikia has entered the search business. They want to be open, transparent and as I learned from within Wikia, they want to be multi lingual in a few months. I will be SO happy if they allow people to tag information on the Internet that is important to them as being in a specific language. Because this will allow us to point to documents that exist on the Internet.

Thanks,
GerardM

Wednesday, January 09, 2008

Just another great day

Today, is a good day. All kinds of things have happened that made me quite happy. One of the best bits of news is the news that Nikerabbit has successfully tested imported "gettext" data into BetaWiki. To understand the relevance, it means that professional translators can use their tools like OmegaT and translate the MediaWiki system messages or the messages that go with MediaWiki extensions off line. This is a major boon specifically for the less resourced languages.
Thanks,
GerardM

Tuesday, January 08, 2008

CentralAuth and FlaggedRevs

CentralAuth and FlaggedRevs are MediaWiki extensions. They are quite special because they are extensions that we are waiting for to be implemented. CentralAuth is better known as SUL or Single User Login and FlaggedRevs is flagged revisions.

There may be all kinds of reasons why they have not gone live yet, I am happy to say that localisation is not one of them as the localisation has been done for many languages. There are few practical things that we can do to make a polite point that we would really really really have these extensions go live. Localisation is such a polite way.

I invite you all to check the statistics and if your language is not there yet goto BetaWiki and help us localise the 59 messages for CentralAuth and the 122 messages for FlaggedRevs.

Please Brion, consider this as well in your planning ...

Thanks,
GerardM

Thursday, January 03, 2008

MediaWiki localisation of Marathi

Marathi is a language spoken by some 68 million people in India. The mr.wikipedia has some 15.195 articles, the mr.wiktionary has 267 articles. I am not aware of other WMF projects. I am not aware of other projects in Marathi outside of the WMF.

73.99% of the most important MediaWiki messages for Marathi are now localised. As not all of the most important messages of MediaWiki, MediaWiki is not really useful out of the box. One of the ways the localisation can be improved, is by importing messages from projects. SPQRobin imported some 500 more messages from the mr.wikipedia and after running several sanitizing scripts, they were committed.

SPQRobin does not speak Marathi but it is probable that he did a great job. It would be good if someone has a look at the Marathi messages. Hopefully not only to check the existing messages, but also to work on the messages that still need to be done.

Importing from a project is one of the obvious methods of enriching the MediaWiki localisation. Proof reading the localisation is important and will improve the quality of the MediaWiki experience. Hopefully someone will come along and do this for Marathi

With the BetaWiki developers actively looking at improving the MediaWiki localisation, allowing MediaWiki to already have five separate forms for a plural, it may be that languages like Marathi or Hindi have their own needs. :)
Thanks,
GerardM

Monday, December 31, 2007

Happy New Year

Some of my friends are already in the New Year, some like me are celebrating the old and awaiting the New Year. The SignWriting you see is the Czech sign language.

One of my hopes is for SignWriting to do well in the new year. I hope that many people will learn to read and to write their sign language. I wish us, that the best is in front of us... Happy New Year :)
Thanks,
GerardM

Linguistic tolerance

The Wikimedia Foundation has a policy about what linguistic entities it accepts and what linguistic entities it does not accept. When a linguistic entity is recognised as such and has an ISO-639-3 code, it is considered a language. As a rule, the language committee gives a conditional approval to requests for languages that have a code.

There are problems with this policy. As I wrote elsewhere, a language may be dead. A dead language is a language that is no longer actively used; there has been no new terminology, nobody is using it actively, good examples are Hittite (hit) or Akkadian (akk). In my opinion they can have a Wikisource but a Wikipedia is problematic because you cannot write in these languages for a modern public without changing the language into something completely different. It does not even make sense to have a MediaWiki localisation for such languages.

There are more problems, what to do with languages where from within the culture it is prohibited to write the language down. What to do with languages where there are few people speaking a language. What to do with languages where few people are truly literate for their language. What to do with constructed languages?

The biggest issue with all these issues is one of competency. Who is competent to judge that a language is truly dead. At what level are there sufficient people in a community to support a language for a WMF project. How do we judge the quality of our projects and as importantly who is to judge? Also does the WMF have a responsibility for less resourced languages.

Brianna blogged about the Volapük wikipedia. For her and for many others, Volapük became an issue because they had the audacity to create enough articles to be noticed. People like Brianna feel offended because it upsets the notion of what Wikipedia is. Brianna introduces the notion of a "language ego" but I am sure she will agree that every non dead language deserves its place under the sun and only the people that communicate in a language are the ones that can realise a WMF project. The good news is there is plenty of sun and it is not expensive to have another language.

When people equate artificial languages with languages without merit, they have a problem. Many languages have started out in this way. One of the more interesting examples is Italian, it was standardised by Dante, used as a lingua franca until the unification of Italy when it became an official language. Another example is Sardinian where a constructed merged linguistic entity that has not been recognised by the ISO-639-3 registrar, is recognised in Italian law.

When you compare Volapük with Klingon, the biggest difference is that Volapük allows you to express all modern subjects.

For me the issue of the Volapük Wikipedia is a non-issue. I know three people that speak Volapük, not all of them are involved in this project. Given the competency of the people that are, there are no issues in getting the information in Wikipedia right. Wikipedia has always allowed for a project to evolve and insisted on the independence of communities. I am sure that in the end both Wikipedia and the Volapük Wikipedia will emerge stronger from all this.

Thanks,
GerardM

Sranang Tongo

Sranan Tongo is a creole language spoken in Suriname. A request was made for a Wikipedia in this language and as it does have an ISO-639-3 code (srn), granting a conditional approval was a formality.

What is becoming a comfortable routine is requesting the great people at BetaWiki to support another language. People that know the language can now start with the most used messages. I wish the proposers for this new project well, and I hope to hear from them when they consider to be ready for the big time :)

Thanks,
GerardM

Saturday, December 29, 2007

Farmer; tactical technology

Farmer is an extension for MediaWiki. Farmer helps with the maintenance of Wiki farms and enables changes to the configuration from a Web interface. The reason why I blog about it is because it is one of those little gems that may make a difference in making MediaWiki more popular.

MediaWiki can be found in an NGO in a box. This is an initiative of Tactical Tech, an organisation that aims to "demystify technology for non profits". The big selling points for MediaWiki are the many people that know MediaWiki through Wikipedia and the many languages that are supported by it.

Recently farmer was welcomed as an extension in BetaWiki, and the developers active at BetaWiki have been working hard at improving the messaging of Farmer. Farmer will make the maintenance of MediaWiki less challenging. With its messages translated in more and more languages, MediaWiki becomes more and more tactical technology.

Thanks,
GerardM

Friday, December 28, 2007

Firefox 3 beta

I have been bold, I have installed Firefox 3 beta. I cannot say it is all good but it is for many things much better. I would not have installed it without Firefox supporting Chatzilla.. Chatzilla is mandatory for me.

What I like:
  • When I click on an Arab text it will cleanly select the whole word.
  • URL's are shown much more cleanly
  • Firefox is still a great program, it seems more stable and responsive
What I do not like
  • I do not like the presentation of the browsing history
  • There is a bug that has a tab go to the beginning of the page
Thanks,
Gerard

Monday, December 24, 2007

Localisation fast and furious

When things reach a certain maturity, visible things can happen really quickly. Another Christmas present is this presentation of the localization of MediaWiki. It shows the quality of the localisation of the MediaWiki software.

You will see a lot of red, not good. This is a typical situation of the "cup being half full" as the list is includes more languages. With a new visualisation you can not compare. So you do not notice the many recently added extensions. You do not notice the many recently added languages. You do not notice that there would have been more red for many languages. :)

If you want to help MediaWiki, help us improve the MediaWiki localisation for your language at BetaWiki..

Thanks,
GerardM

Friday, December 21, 2007

BetaWiki exports .po files

BetaWiki has given us a splendid Christmas gift. Nikerabbit, Hashar and Siebrand have developed an export tool for the MediaWiki system messages into the .po format. This format is the standard used for many open source applications.

The most important thing about this format is that there are many tools that support it for off line translation. Many translators will not work on-line. For many languages we do not need to accommodate off line translation when there are sufficient people willing to maintain the localisation. However, when you look at the statistics, you will find that there are many languages not supported or poorly supported in MediaWiki. Hindi is a good example. The Wikipedia is well localised but for Hindi only 3.99% of the system messages is translated. Hindi is spoken by 180.000.000 people...

For Hindi .po files will not be the solution. Collaboration between Indian Wikimedians and the BetaWiki administrators will be a better solution. There are also languages like Neapolitan where it helps to localise and then have the localisation proof read. When the number of collaborators for a language is small, it is typically easy and safe to work off line. You do not have to wait for loading and saving, you can combine it with a translation memory and make it efficient.

Nikerabbit has now created a .po importer. What he is looking for are translations to test this new functionality...

Thanks,
GerardM

Thursday, December 20, 2007

A great Christmas card


The question: what language and, what does it say..

Happy holidays,
GerardM

Wednesday, December 19, 2007

WOSI, a really cool Open Source/Standards project

WOSI is a project of a Dutch school, the HVA, working on a software environment for "woningcorporaties". A woningcorporatie is an organisation that is involved in public housing. As an organisation they are genuinely capital intensive, their IT requirements are complex and evolving.

The objective of the school is to provide an environment for their students that will give them a real feel for what it is like to work in the ICT business and teach them Open Source and Open Standards. Students that are part of this project will experience that a project does not start from scratch, there is always something to build upon, there are always conflicting requirements, there is always the need to ensure interoperability because the use of Open Standards is a precondition.

The WOSI project is in its second year and, it is growing in size. Students from other disciplines are getting involved as there is overlap with other specialties like communications and marketing. More woningbouwcorporaties are interested in the project as well as other schools. The great thing is that as Open Source and Open Standards are key to this curriculum, professionals will be released to the job market that know how to apply these notions in real world scenarios. This is likely to prove the biggest boon to this really cool project.

Thanks,
GerardM

Tuesday, December 18, 2007

Inter operability is important

Wikipedia and particularly the English language Wikipedia is a rich resource of information. The amount of information in it is staggering. Much of the information is duplicated in other Wikipedias and other websites. This is great. Because with more applications for the same data, more eye balls will find what is in error.

I am subscribed to the DBpedia mailing list and today I read about errors in Wikipedia that had to do with Wigan and Manchester City. Errors were found and the gentleman wrote that he can and will make the necessary updates. His question is when will the DBpedia reflect the changes.

When the data of Wikipedia is analysed with tools, and when the results are found to be of value, it adds relevance to what enables this collaboration. It typically relies on the availability of dumps. When the data is analysed, a new work emerges. When it has a completely different format, it is possible to mesh it with other data sources. This in turn will help establish the validity of the Wikipedia data and will allow for the extension of the data.

When multiple data sources are meshed, the issue of copyright and license raise their ugly head. You can create static and dynamic meshes. In a dynamic mesh you can build the mesh depending on what the person has access to. In a static mesh you can only include the data that is still available to the least privileged person who will get access to the data.

The consequence is that many people, organisations will mesh sources, manipulate data, publish and not indicate what all the sources are. They will not do that because they do not want to be bound by all kinds of licenses and because they do not want to be hassled.

This DBpedia example shows that the presentation of facts is important. It demonstrates that interoperability will result in a better Wikipedia. It is important for Wikipedia to be as open and engaging as it can be. Frankly, when people analyse our data in a similar way to DBpedia, it is a new work it should not be considered derivative. Best practice is to publish sources and this, more then the viral nature of a license like the GFDL or CC-by-sa, will drive collaboration and give Wikipedia more relevance.

Thanks,
GerardM

Sunday, December 16, 2007

Localisation of MediaWiki

When you wonder in what languages MediaWiki has been localised, and to what extend the localisation is usable, BetaWiki has some great statistics.

The Localisation statics show the languages that have a central localisation and the percentage of the messages that have been done. It clearly shows that the MediaWiki localisation leaves a lot to be desired; for some 144 languages less then half of the messages have been localised. At this moment there are 235 languages known to MediaWiki. When you compare this to the 253 language that have a Wikipedia and add the languages that are starting in the Incubator, you get a clear picture of how much effort is needed to better support the readers of MediaWiki projects.

When you look at the statistics, the glass is half full, and it is filling. On average five languages are introduced every month and more then 500 messages are translated every day. The languages that are in the Incubator are doing well Seeltersk for instance has done an astonishing 99,3%.

One of the latest innovations in the BetaWiki are the core top 500 messages, they contain the most important messages and with these messages translated, MediaWiki is usable for a language. BetaWiki has a dedicated team of people that make MediaWiki and as a consequence MediaWiki projects usable for many people. With your help, we can improve the localisation even further. One message at a time will slowly but surely provide proper support for all the languages MediaWiki supports.

Thanks,
GerardM

Friday, December 14, 2007

It is perfect after all

In my latest post I wrote about the Oostvaardersplassen, today I received a mail telling me that fish will in future be able to swim into and out of the Oostvaardersplassen.

This makes me perfectly happy. Now I know how the water flows from the Oostvaardersplassen into the "Wilgenbos" and I know that there will be a lot of work that needs doing. But when Staatsbosbeheer, as it does, states that fish will be able to swim in and out .. really great news.

Thanks,
GerardM

Wednesday, December 12, 2007

Vindication of a kind

One of the Wikipedia articles I am proud of is the Dutch article about the Oostvaardersplassen. The Oostvaardersplassen are close to where I live and I think it is one of the best examples that nature is something that not only evolves, but also can be engineered. I have followed its development with considerable interest and my favourite point of view has been that the water management has been detrimental to the natural diversity.

I visited the Oostvaardersplassen this weekend, and I learned that a small dyke will be removed leading to a more natural distribution of water and a more dynamic water level. This will have a huge impact on the fish stock; the current population of mainly mature carps will make room for many more smaller fish. This will allow many small herons and other fish eaters finding their niche.

The one remaining question for me is if fish will be able to freely migrate in and out of the nature reserve. It would be grand if this is the case.. As only one dyke is mentioned, I do expect it to be great but not "perfect".
Thanks,
GerardM

Monday, December 10, 2007

Burglary

The word of the day for OmegaWiki should be burglary. It is not as another word had already been entered. I was sleeping and woke up because I heard the breaking of glass. I looked out of my window and saw someone breaking and entering. I called the police, they arrived quickly..

After all the excitement, I find it hard to go back to sleep.. Anyway, this is real life drama. Not dramatic, but it keeps me from sleeping.

NB the word of the day is íshokkí.

Thanks,
GerardM

Wednesday, December 05, 2007

Shameless plug ...

A friend of mine send me what she called a "shameless plug". I agree with her, there is no shame in announcing that wikiHow is supporting the Dutch language.

As I absolutely approve of great projects doing great things, I am happy to shamelessly plug wikiHow and I wish it and all its language versions great editors and a great audience.

Thanks,
Gerard

Monday, November 26, 2007

Wiktionary upset

When I look at the Wiktionary website at the moment, it does not show yet that the French language Wiktionary has more articles then the English language Wiktionary. I think it is absolutely wonderful because if anything it shows that you should not take things for granted in Wikis.

Not taking things for granted is a healthy attitude. There is an inherent bias against the French language Wiktionary in the Alexa numbers. However, I am impressed by the numbers quoted.

All the bigger Wiktionary projects have used bots to build up their content. I can imagine that a healthy rivalry will make the numbers go even higher, this would benefit the users of Wiktionary because I trust the Wiktionary communities to watch the quality of the content :)

Thanks,
GerardM

Wednesday, November 21, 2007

Pride in a language


Many people feel strongly about their culture, their language. What I find special are the people that go the extra mile to promote their language. It is therefore that I am grateful when I notice languages like Spanish, Georgian, Breton and now Eastern Yiddish having a champion that make a difference.

It is especially interesting to see how with an ever increasing amount of terminology, the information becomes rich. Rich both for the people who are interesting in learning the language and also for the people that want to learn other languages starting from these languages.

I am grateful when people find in OmegaWiki a tool that helps to document and service their language. It is an imperfect tool but its redeeming quality is that it is getting better as we go along.

Thanks,
GerardM

Tuesday, November 20, 2007

Thank you Mycom

Today started horribly; my Skype was not working and my headset was to blame. I could not call using skype so I had to use my plain old telephone to call abroad :( I then tested my system and it was indicated that my microphone was not working. I went to my computer supplier and, I was told that my headset came with a two year warranty !!

I got home and it was still not working. So I played with my configuration and I could not get it to work. So I went again to my computer supplier Mycom and we fiddled with all kinds of values. For whatever reason we got it to work but we were not able to pinpoint what the issue was. I think that it may have to do with an upgrade of the Skype software that I did the other day but I am not sure. In the end, the friendly service I got was the highlight of the day :)

Thanks,
GerardM

Friday, November 16, 2007

69 page document in SignWriting

I received an e-mail that mentions a 69 page document is SignWriting. At this moment a document in SignWriting is still considered to be exceptionally large. It does prove that people are able to write documents in their sign language that are this big. As more texts are written it will be not be special much longer.

The reason for me to mention it is that it indicates that SignWriting is stepping over a threshold. It is enabling American Sign Language to be a literary language. This is another step closer to the realisation of a Wikipedia for ASL.

Thanks,
GerardM

Tuvin‎

Tuvin is a language spoken in Russia, China and Mongolia. There is a wikipedia project on the Incubator that is dormant. What is exciting is that there is activity to localise MediaWiki in the Tuvin language.

On the Tyawiki there is a MediaWiki installation that aims to create a repository about Tyva. Withthe localisation of Tuvin in MediaWiki, they will be able to do a much better job.

I welcome this first localisation effort I am aware of that is driven from outside the Wikimedia Foundation. It demonstrates how MediaWiki is getting recognition for the outstanding software it is.

Thanks,
GerardM

Thursday, November 15, 2007

Flags for languages

I got into a discussion about what flag should be shown for a particular language. There were two groups of people one side was in favour of a historical faction and the other was in favour of another historical faction. The argument was quite heated. At some stage I asked what the flag for English should be ...

Languages are spoken on both sides of a border. Languages are spoken by people who do not recognise countries or flags. Languages are separate from nationhood. Consequently it is in my honest opinion wrong to associate languages with flags. There are no obvious symbols for languages and for many languages they would share the same flag. OmegaWiki is not likely to ever associate a language with a flag.

Thanks,
GerardM

Statistics

On Alexa, OmegaWiki has for the first time gone through the 100.000 traffic rank. This number gives in Alexa better graphics so it is really welcome news. The daily rank of 99.813 is extraordinarily compared with our three monthly rank of 485.925. It however indicates that something is working in our favour or maybe we are doing something right.

When you compare the OmegaWiki statistics with the Wiktionary stats they are doing really great. With a daily rank of 1.601 it is time for the Wikimedia Foundation to demonstrate that they have a valuable resource in Wiktionary :)

Thanks,
GerardM

Tuesday, November 13, 2007

Political science

With some dismay I read this article on the BBC-website. It is about science and the cost of science. It is said that there should be better planning so that ahead of time it is clear how much a given project will cost. This should mean that "the wider scientific community and industry should contribute to the decision-making process".

Effectively the right honourable gentlemen is looking for a reduction in the cost of science by increasing the involvement of even more people to select and manage scientific projects. Effectively this will lead to less money for doing science. Effectively it will not be the scientists who select the projects that are considered to be of scientific value.

The dismay I feel is because it is likely to lead to more yet "politically correct" science. Science that has more to do with what the expedient results should be and not with scientifically relevance. When you consider the huge amounts of administrative and other overhead it is a wonder that scientific research is still practised.

Thanks,
GerardM

Sunday, November 11, 2007

When is a project alive

A month ago I blogged about dbpedia. Dbpedia is very much alive. I have had a look at it and like what they do. I subscribed to their mailing list and again, dbpedia is very much an active project. As always there are things I do not like; their ideas on copyright and licenses are defensive and as a consequence overly restrictive. It prevents cooperation in stead of fostering cooperation.

Today, an anonymous person replied to this blog entry. The suggestion is made to cooperate with the SWAD Europe group. They have a website, a blog but it all stopped in 2004. So I am wondering about all these projects, all this effort that just stops. Projects that may be valuable and given that people promote it in 2007, may still be alive. For me there is no way of knowing.

I have an idea how I would use semantic data in OmegaWiki. What I am not so sure about is how semantic web applications would use OmegaWiki data. In essence OmegaWiki is multi-lingual and exporting it in anything but a machine readable version only, would strip what I think is valuable in OmegaWiki.

Collaborating with for instance a SWAD Europe group makes sense. People can suggest cooperation, it should however be a two way street. Just pointing that there are others does not help me much.

Thanks,
      GerardM

Wednesday, November 07, 2007

Congratulations OLPC

I was happy to read this..
Thanks,
GerardM

The World Language Documentation blog


It is with pleasure that I inform you that the World language Documentation Centre has a blog. As the WLDC is ambitious in what it wants to achieve, and as many of these objectives will take time, it is great that there is a blog where the board members of the WLDC can publish about they find of relevance.

I hope that the many members of the WLDC board will find the time to blog because this will help you appreciate the amazing qualities that you find in these people. As I have the privilege to be on this board as well, some of the subjects that I have written about in the past will now be covered on the WLDC blog ..

I hope you will find the WLDC blog of interest to follow it in your RSS reader.. :)

Thanks,
GerardM

Sunday, November 04, 2007

Stuttering

"A speech disorder in which the flow of speech is disrupted by involuntary repetitions and prolongations of sounds, syllables, words or phrases, and involuntary silent pauses or blocks in which the stutterer is unable to produce sounds". This is as I have always understood stuttering or stammering as the Brits have it.

With great surprise I learned that stuttering also occurs in sign languages. I learned this from a mailing list that deals with sign languages and SignWriting. The implications are quite profound. It means that stuttering is not necessarily a speech disorder and consequently when it is not, speech therapy does not work.

This similarity in the problems between signed and spoken languages indicate that the format of communication is incidental. The same mechanisms are at play and therefore one is as good as the other. To me this seems obvious many people rate their own method of communication as superior. The spoken language is superior for when communication with me as I am dumb when it comes to signed languages...

Thanks
Gerard

Saturday, November 03, 2007

Who is Frank Thompson, and why include them in Wikipedia

Frank Thompson features on the "List of mayors of Yarra". He is "blue linked" so there must be an article on him right? Clicking on the link gets me Frank Thompson who was a member of the house of Representatives. It is more or less easy to fix and I am sure that someone will.

For me it is interesting to see how Wikipedia deals with what is considered relevant. To me these people are irrelevant, both misters Thompson are no longer in office. But they are deemed to be noteworthy enough to link to where might be an article.

In a similar way there are articles about pop stars who had a single hit in 1962, there are articles about wide receivers that only played one season.. There is a lot of information that is of no importance and that is fine.

What astounds me is that when an article is written in the German and the English Wikipedia about Kotava, a constructed language, it is speedily deleted. When it is then indicated that this language is on route to be recognised in the ISO-639-3 code, the comment is speedy deleted. The article that was deleted was more then a stub, it cited sources and I did not write it.

I would love to understand why a mayor of Yarra, a 1962 pop star or a 1956 wide receiver are "relevant" and a language like Kotava is not.

Thanks,
GerardM

Tuesday, October 30, 2007

Kotava - another constructed language

Kotava is a constructed language. There is yet not article in the English Wikipedia, there is an article in eight other language.. (myth busting; if it is not in the English Wikipedia, it is not in Wikipedia).

At this moment Kotava is not eligible for a Wikipedia, it will not be enabled for editing in OmegaWiki. It does not have an ISO-639-3 code yet. What is special is that there are clear indications that the process for a code is under way. The code is likely to be "avk".

At OmegaWiki there is a Kotava enthusiast who has started a lot of the preparations for another language. It will be given once the code is official. For a Wikipedia, they may ask for a Kotava Wikipedia. With the ISO process under way, the language committee does not have to do anything until the code is granted.

It is great that the language committee has reserved the right to do nothing..

Thanks,
GerardM

WCN, the network - an unsung hero

When you are at a conference and the networking just works, you will not hear anyone about it. At the Wikimedia Conferentie Nederland NOBODY mentioned the network; it was just there and it just did what it was supposed to do.

GREAT :)

Thanks,
GerardM

Sunday, October 28, 2007

Google docs - published presentation


When I am not signed on in Google docs, blogger, any Google application and I look at the URL that I used in my blog I get this screen. At the very bottom it indicates that I can have a look at the presentation. It then works for me..

Thanks,
GerardM

Saturday, October 27, 2007

WCN

Today the Wikimedia Conferentie NL was held in Amsterdam. It was a great occasion. It was impossible to do justice to the program; they had three tracks and to chose one presentation over the other was an injustice to what was missed.

I had the privilege to give a presentation, and as I had to do some serious travelling to be there, I considered to what extend I could reduce what I took with me. I decided that with the latest Google application I did not need to bring anything. I could rely on there being a network, Kim brought his Merakis who are still on the old functional software and assuming one functional lap top should not be a problem either.

To prove that it was indeed the Google presentation tool. I selected one of the available backgrounds. It is different. What I need in a presentation is basic. I need to show some texts and some screen dumps. I think there is some need to polish the handling of the screen dumps.

One of the nice things is that the presentation can be made available. So have a look and let me know what you think.

Thanks,
GerardM

Monday, October 22, 2007

Breton

Breton is is a Celtic language spoken by some of the inhabitants of Brittany (Breizh) in France. According to Ethnologue over half a million people speak the language.

There is an organisation that is actively promoting the Breton language. I am really happy to inform you that they have taken the trouble to localise the system messages of OmegaWiki and have started to add translations in Breton. I have send a file with many of the languages that are in the ISO 639-1. This combination will localise most of OmegaWiki in Breton.

The argument that proved convincing is that you get all the information we have. By steadily increasing the translations available in Breton, the experience will improve.

Thanks,
GerardM

Saturday, October 20, 2007

Import, export

In OmegaWiki we provide some statistics. One of them is a breakdown of the Expressions per language. It is interesting because it shows what people are working on. It is a good indicator because after the initial import from GEMET, all the new work was done by hand. We are now experimenting with the import and export of data that are in collections and this will change things quite a bit.

The export creates a txt file with the columns separated by tabs. The columns are the number identifying the DefinedMeaning, and combinations of the Expressions and Definitions. The reason why we start with this export is because it is still much quicker to translate in a spreadsheet then it is to translate on the web. These experiments are done in a test environment and, we hope to bring it life soon. I can already send you a file when you are interested :)

When we start to import, it is likely that languages can make quite a jump in the statistics. I hope it will encourage people to help us by providing translations particularly for the less resourced languages.

Thanks,
GerardM

Friday, October 19, 2007

Dbpedia

I chatted with Duesentrieb the other day. He mentioned dbpedia. I had another look and it is really a great resource. For those that do not know, dbpedia is a community effort to extract structured information from Wikipedia and to make this information available on the Web.

It does a great job and it is different from that other project that deals with structured information, Semantic MediaWiki, in that it does operate by data mining information from Wikipedia while Semantic MediaWiki is an integral part of a MediaWiki project.

The great thing of dbpedia is that it explicitly encourages interlinking. With interlinking data, data that can be found in another resource, becomes available limited by the quality of the interface.

You might ask how OmegaWiki fits into all this. Dbpedia's information is in English while OmegaWiki allows for the representation of information in any language . Both OmegaWiki and dbpedia link to Wikipedia articles and consequently where the two share a link, the information can be mashed together.

In Wikiprotein there is a large amount of medical information available. This information does link to external databases. Given that the medical articles relate to external databases, there is an opportunity to link the data to Wikipedia articles. With Wikipedia articles linked in this way, more specialists will find their way to the Wikipedia medical articles and this in turn will make more enriched information available.

The thing to consider now is how OmegaWiki can benefit from dbpedia.. one of the issues is the difference in license. Then again, dbpedia provides its algorithms and consequently the result is not necessarily the license that dbpedia posts.

Thanks,
GerardM

Friday, October 12, 2007

Domain names in other scripts

The Washington Post is reporting that ICANN is experimenting with URLs in other scripts. This is good news. For people that do not read or write in a language that uses the Latin script, it is a big handicap to have to type for instance http://ar.wikipedia.org also http://ar.ويكيبيديا.org is problematic. In order to enable people, you have to have the whole string in the appropriate script.

From a localisation point of view, this is one of the ultimate challenges. Consider; the .org has two components, the dot (.) and the org. The org needs an equivalent in all scripts. This top level domain or TLD is just one of many. The ISO 15924 defines many scripts and, a particular combination that is auspicious in one language can be the equivalent of wtf in another. Choosing these codes is not trivial. There are not only TLDs but also ccTLDs or country code top level domains.

We have agreed that ويكيبيدي is Arabic for Wikipedia .. Wikipedia is very much an international movement. We have agreed that Wikipedia is to be used for the Latin script in our domain name. The question is, if ويكيبيدي will be accepted to represent Wikipedia in the Arab script and, when there are multiple ways of writing Wikipedia, do we need to register for all these domains?

To make it even more confusing, what will the rules be when it comes to domain squatting. I can imagine that a brand is only registered for one script and not necessarily for another. I wonder what the position of the WMF will be; I am sure that it has not been considered yet.

ICANN is courageous, they are now experimenting with the technical issues and this will show that it can be done. I expect that the next part will be a proposal on how all the top level domain names are to be "transscripted". Then, it will become interesting because from that moment onwards the Internet will be truly global in its reach and no longer centered on one language or script.

Thanks,
GerardM

Saturday, October 06, 2007

A picture paints a thousand words

I mentioned that in OmegaWiki we are really moving on the localisation. I am thrilled to announce that a lot of work has been done for Spanish, French, Portuguese, German and Dutch. But the icing on the cake is that for two other scripts a start has been made; for Serbian and Georgian the localisation has started.

I think the Georgian script is pretty :)

Thanks,
GerardM

Now with grammatical gender

In Omegawiki we now have support for grammatical gender. When we now how to say it in "your" language, we can show it. :)

Thanks,
GerardM


OmegaWiki vs Semantic MediaWiki

Both OmegaWiki and Semantic MediaWiki are providing semantic support. There are people that have expressed that OmegaWiki should not include particular types of data because Semantic MediaWiki does a better job.

Both OW and SMW are extensions to MediaWiki, so at first face it seems like a reasonable suggestion. The two extensions however do completely different things.

Semantic MediaWiki will shine when it becomes part of a project like the English Wikipedia; when key data that is in the article is marked, it will provide a great improvement in making these facts available. SMW even provides a really rich environment to query the information. It is absolutely great and it is absolutely mono-lingual.

OmegaWiki is at this moment very much a stand alone application. It does not derive data from anything, it is great at presenting the same data in many languages. This means that when we know that the information exists, we will show it in "your" language.

SMW can export and when OW can import, we have the best of both worlds; that is to say we have the best of both worlds when we can link the resources to each other. As OmegaWiki is not encyclopaedic and does not want to be, it is our stated intention to link to Wikipedia articles. As the SMW is tightly linked to the Wikipedia articles, this may be just the trick.

The suggestion that OmegaWiki would only have the semantic information that is in Wikipedia is incorrect. In Wikiprotein we already have a real rich set of annotations of proteins. This is the kind of information that is not encyclopaedic. It enables scientists to maintain information on "their" proteins. The language of the science of proteins is English, however many are known in other languages. It is rich that all this can integrate as it brings diverse information together.

Both OmegaWiki and Semantic MediaWiki have their own, and different strengths. Within Open Progress we use Semantic MediaWiki for our internal wiki. It works absolutely fabulous. I love both MediaWiki extensions!

Thanks,
GerardM

Thursday, October 04, 2007

Changes in the user interface

A DefinedMeaning in OmegaWiki can have a lot of data associated with it. China borders on so many countries and seas that it is just a bit much. So it makes sense to bundle certain types of information together and keep them separately. This has a profound impact on what the data looks like. So far I was happy when we had the data on the screen but now it becomes possible to think where does it makes sense to have the data. I asked Erik if the incoming messages could be at the bottom of the page..

Well some things seem like miracles and we can have them in five minutes.. The problem with the collations is that at this moment it has to be by hand. This makes that it does not scale. With the "borders on" example however, we have a great showcase WHY we need to be able to sort these texts and also why they should be Expressions in OmegaWiki like all the other information that we hold.

Really, I could not be more happy with the progress that started to happen.. :)

Thanks,
GerardM

Wednesday, October 03, 2007

Is it a Wiki ?

OmegaWiki is a wiki. Some people however disagree; they consider the fixed format conclusive evidence why it is not. As the software is maturing, this argument loses a lot of its lustre. Obviously we have been saying that these people are wrong all along :)

The best argument why OmegaWiki is a wiki is because people can add/ change little items one at a time, they are not compelled to do everything at one go. As this is acceptable, our data has to be correct but does not necessarily need to be complete.

With the new terminological support; we can now indicate that a language is a language, all kinds of additional information can be added once you have stated that it is a language. Have a look at French for instance, the "incoming relations" and the annotations provide a lot of information. As the moment of writing it is not yet clear that French is an official language of France for instance.

With the expansion of existing classes, with the expansion of the class attributes information will become more available and integrated. The best bit is that as the OmegaWiki specific user interface can be translated in many language people will be challenged to ensure that the right terminology is used in translation.

With the new functionality OmegaWiki became much more wiki. It is there for every one to see, and everyone is cordially invited to have a look, create a user, add some Babel templates and have a go at it.

Thanks,
Gerard

Monday, September 24, 2007

Self promotions on BBC-News

The BBC-news website is my primary website for news. It provides great information and I have read it for many years. As I am driving less nowadays, I do not listen as often as I used to to the BBC-Worldservice (648 AM).

Recently there has been an increase in the number of video fragments. At the same time I find that I am less likely to watch them. The reason; every fragments is now preceded with a promotion and it is annoying and distracting to the point where I do not bother any more.

Thanks,
GerardM

Sunday, September 16, 2007

Zuckertüte

When kids go to school for the very first time in Germany, they get a " Zuckertüte". A Zuckertüte is a big carton cornet filled with sweets and little gifts.

The word Zuckertüte is a good example of a word where I do not expect translations in any other language as it is a real German tradition. This does not mean that there cannot be translations of the definition.

I hope that the grandchildren of this little lady have a great day at school tomorrow :)

Thanks,
GerardM

Friday, September 14, 2007

Shtooka

When I started to record pronunciations for Wiktionary, I used Audacity. It is a great tool but hey when you are a serial recorder, you want to have a tool that assist and helps you to do it it efficiently. Shtooka is a great tool. I think I blogged it in the past.

Now they have surpassed themselves. There is now the Shtooka explorer. It allows you to listen to the many, many pronunciations in several languages they have recorded. It is an absolutely gorgeous application that really shows off an already great project :)
Thanks,
Gerard

Thursday, September 13, 2007

Social networks

I have been using social networks for some time now, LinkedIn, Plaxo, ecademy and facebook is where you can find me. As I have invested a considerable amount of time, it is relevant to consider if they are worth the time and effort.

LinkedIn, Plaxo and ecademy are at a considerable disadvantage because in order to get the full functionality, you have to spend money. It then becomes relevant to understand the potential benefits for these networks. All three provide basic functionality that is for free and as there is no real overlap yet, I maintain a presence there.

Facebook is what I have been looking at lately. What it has right is that it tries to connect people, the groups they belong to, the causes they champion and the organisations they are associated with. When you combine it with the potential to program extra functionality for facebook, you get a more compelling package then its competition.

It is however lacking in other ways. For one the way it has its security is minimal. I do not mind to tell the world that I am on facebook, but I prefer that only friends and friends of friends can see who my friends are. I would happily leave Plaxo behind when facebook had the same security for sharing personal and work information.

The one thing all the social networks have in common is that they are proprietary. To make their functionality useful to me, I have to trust it with my information. I do however not know if the implementation of their security can be trusted. Were they to use something like A-Select I would feel more comfortable because it would allow enough eye balls to vouch for the authentication process. Even though it would be a great step forward, it would be better if all the software were Open / Free software. Authentication is one aspect, the authorisation within the application itself can still make my data insecure.

Who to trust, why to trust ...

Thanks,
GerardM

Tuesday, September 11, 2007

Hehe

Hehe is a language spoken in the Iringa Region, south of Gogo in Tanzania by some 750.000 people. There is no article about this language on the English Wikipedia but there is an article on the people who speak this language; the Hehe.

Even though there are many references at the back of the article, it is marked as not citing references or sources. I do agree however that this article would benefit a lot from wikifying and the creation of supporting articles. It is a good example of the amount of work that needs doing to make the English Wikipedia relevant as a resource for Africa.

Thanks,
GerardM

Monday, September 10, 2007

Sassarese and Sardinian

There is a Wikipedia in the Sardinian language. It uses the sc ISO-639-1 code. What was known as Sardinian became srd in the ISO-639-2. In the ISO-639-3 it was recognised as a macrolanguage; practically what was called Sardinian was split into four languages.

The Italian government has officially recognised the Sardinian language or the "Limba Sarda Comune". This is in essence a constructed language as it tries to make one language out of the four "dialects". One of the effects has been that some people prevent others from writing in one of the four languages on the sc.wikpedia.

The language committee of the Wikimedia Foundation has a request to approve a new language; one of the Sardinian languages, Sassarese with ISO code sdc.

There are two problems to deal with:
  • The "Limba Sarda Comune" is not recognised as a language
  • The proponents of the "Limba Sarda Comune" reserve the sc.wikipedia for their language
This issue is political. The first thing that I understand when you go to the official website is the notion of identity and indeed, to create one Sardinian identity it would be instrumental to have a unifying language. However, the map of the Sardinian languages is clear, the island is divided in four.

Given that the language committee has as one of its rules that political arguments are not accepted, there are a few conclusions that we should make.
  1. Sassarese can have a conditional approval
  2. We urge the proponents of the Limba Sarda Comune to ask for the recognition of this newly constructed language from ISO.
I have had a chat with Debbie Garside about all this, and I understand that it is necessary to apply for an ISO-639-3 code before an IANA language code is likely to be approved. At least fifty published works in the Limba Sarda Comune will be required.
Thanks,
GerardM

Saturday, September 08, 2007

Wikizine but more relevant WalterBE

Walter is a long-standing member of the Wikimedia Foundation. He is a steward, he is the editor-in-chief of Wikizine. He organised the first official meeting of people of the nl.wikipedia .. there were only two people and we had fun.

People who know Walter know him as a soft spoken can-do person. He has done much of the organisational work on the Dutch Wikipedia, organising elections, being involved in OTRS from the start. He is the press contact for Belgium ... As a steward he did much good, he is a member of the communication committee ...

Walter has indicated that he has grown away from the community and as such his motivation for Wikipedia and Wikimedia stuff has gone downhill. He does not feel that he can properly represent the WMF and the community, he is disappointed in the lack of cooperation around Wikizine ... What has prevented him so far to stop is his sense of responsibility to the Wikizine readers. Wikizine has been a labour of love for Walter, there has been little input from the community to inform Walter about the latest, no people except for proof readers who shared the burden of this well received periodical.

Well, I can only be sad that Walter finds his commitments a burden. I do hope that he will know and remember how much he is appreciated for the work that he does and has done. Really, to me Walter is one of the most important Wikimedians.

Thanks,
GerardM

Friday, September 07, 2007

My friend Bèrto 'd Sèra

I think Facebook sucks, they are not able to write my friend's name correctly ...
Thanks,
GerardM

Kamusi, The Internet Living Swahili Dictionary has been taken offline.

Kamusi is a great lexical resource for Swahili. With a lot of great effort this became one of the really relevant Internet projects of Yale University. It is with sadness that I found that it has been taken off line.

It is sad, that such a sterling effort is endangered on what seems to me a minor issue. It is sad because the many Kamusi's users are now without their Swahili dictionary.

I contacted Martin Benjamin, the editor of Kamusi and, I learned that the World Language Documentation Centre is willing to help out with the hosting of Kamusi. Martin and the WLDC are looking for the best way forward; maybe another university can take over where Yale has dropped the ball, maybe Yale will reconsider ...

When I know how to get to the new Kamusi website, I will let you know.

Thanks,
GerardM

Monday, September 03, 2007

Cool application

I found this tool called Touchgraph Google Browser. It allows you to see how a particular website is connected to other websites. It gives you a nice presentation with links between the many websites that are connected. It is nice to compare for instance a wikipedia.org, a citizendium.org and an omegawiki.org.

Thanks,
GerardM

Saturday, September 01, 2007

A computer that works for Luna and Marco

Luna and Marco live in Italy. In Italy it can get hot. It gets so hot that the computer they use overheats. Their computer only works early in the morning and late in the evening. They can push at the edges a bit because Marco found that an Ubuntu live CD gives them some more time then Windows XP does.

Their mother, Luna and Marca are five years old, also has a computer. Her computer only works reliably in the hot Italian summer with an external fan pointing at the computer. The computer is raised a bit from the desk to improve ventilation even further.

In all the talk about computers something simple like the environment is hardly mentioned. PC's and laptops are thought to work in an office environment. In an office environment it is expected that the temperature is regulated.

The OLPC is made for kids and it will work in a hot environment. Luna and Marco would love to have one. It is sad that their mother has to use a computer that cannot take the heat and, a computer that is not as cool.

Thanks,
GerardM