Monday, April 23, 2007

Lies, damned lies and statistics

It has often been said that the only "reliable" statistics are the ones that you compile yourself. Strike that, I am not great at this science of statistics, I just use them to blind people with what is made obvious in this way. A week ago we celebrated that we had 15.000 Expressions for German. Today we celebrate that we have over 10.000 Expressions in Italian.

According to Kipcool, the statistics are wrong. His statistics take into account the fact that we occasionally have reason to delete Expressions. According to his figures we have 14.020 German and 9.776 Italian Expressions. This means that we may celebrate the round numbers again at a later date.. :)

For the Destinazione Italia project in OmegaWiki, we will need translations in many languages with an emphasis on European languages. As we already have a lot of content in Italian, it is important to mark what is already there also is known to be part of this collection. In this way do we know how many Expressions are there to process.

It seems to me that numbers, statistics tell a story, they tell as much about what is represented as about the person who uses the data. It is after the collaboration on statistics that you find what the numbers actually mean and what they truly represent. After such a process, they have become the statistics of a project and are owned by its community.

Thanks,
GerardM

Saturday, April 21, 2007

Wiktionary quality issues

On the Wiktionary project I run the interwiki bot. The process is simple; when an article exists in another language spelled exactly the same, I create an "interwiki" link. This allows you to see the information on another language Wiktionary. This process is an automated process, it works on all Wiktionaries and it is an unattended process.

I have received a request from the Polish Wiktionary to stop adding interwiki links for the Russian and for the Vietnamese Wiktionary. The reason given is one of quality. On the Russian Wiktionary many of the articles are created by a bot and they do not provide good information. An example is dispersion, there is nothing really in there. The Vietnamese Wiktionary is more problematic because a bot was used to generate declension and conjugation tables of Russian words and they got it wrong.

The Russian Wiktionary has some 81.000 empty shells and refuse to remove it. The Vietnamese are not willing to remove there incorrect data.

I have been asked to stop including the Russian Wiktionary and the Vietnamese Wiktionary when I run the interwiki process. To be honest, I run the bot as a service and I do not think it is the right thing to do. I think the Vietnamese are wrong not to correct the wrong data that they have. I am less sure about the Russian approach; in essence it is a stub. However, creating a Wiktionary in this way is like stamp collecting; you can look at it but there is not information about it.

Given how the process works, I am not sure that I can exclude either the Russian or the Vietnamese Wiktionary. The way it works is that I run explicitly on all Wiktionaries. When I exclude Russian or Vietnamese, I will probably end up removing all references to these projects. They are the third and fourth Wiktionary is size.

When I do not exclude the Russian and the Vietnamese Wiktionary, the bot may end up being blocked on the Polish Wiktionary. This will also kill off the interwiki process.

From my point of view, using bots to generate content in a Wiktionary only makes sense when there is at least a link to the word in the base language. When the initial creation of stubs is followed by the enrichment of these stubs it is acceptable. For having information that is completely wrong, there is no excuse.

The question is, will there be a discussion about acceptable practices in Wiktionary. The question are:
  • Can the Polish demand what they do?
  • Is having a project that consists mainly of stubs acceptable?
  • Is having incorrect data acceptable?

Thanks,
GerardM

Tuesday, April 17, 2007

The Uyghur‎ user interface

Uyghur‎ is a Turkic language spoken by the Uyghur people in Xinjiang. According to the English Wikipedia, the language is written in the Arabic, Roman and again the Arabic script. Cyrillic is actively used as well.

The officials script in China for this language is Arabic and has been since 1983. The ug.wikipedia has a Latin script user interface. The ug.wiktionary however is right to left. This means that there is at least a need for a user interface that is either in Arab and one in the Latin script. Given that Cyrillic is another actively used script, we need three message files for Uyghur.

When a language is expressed in so many ways, the current MediaWiki software does not allow for supporting Uyghur in a meaningful way. It will be good when Multilingual MediaWiki becomes a reality. The good news is that the prospects for this are really good. The last status report I heard was that some install routines have to be written and then people can start experimenting with this new functionality.

PS Check out what MLMW offers :)

Thanks,
GerardM


Friday, April 13, 2007

SignWriting

SignWriting is one way of expressing signed languages. There are several ways of doing this, signwriting is the one favoured by most people who actually write down the signed languages. There is a request for a Wikipedia for the American Sign Language or ASL expressed in SignWriting.

I am in favour of such a project. There are however a few relevant issues
  • SignWriting is currently not supported in UTF-8 and MediaWiki expects UTF-8.
  • There is software specific to SignWriting, it can be used for a Wikipedia
  • Due to the technical issues, there is no way it can use the Incubator.
Given the objective of the Wikimedia Foundation, supporting this project is very much something that we should support. ASL is a very different language. Given that the organisation behind SignWriting is very much in favour of the creation of a Wikipedia, there is ample scope for the Wikimedia Foundation and the SignWriting organisation to work together and overcome all the relevant issues.

For me it is important that with this Wikipedia project the culture of the deaf will get a boost. When SignWriting becomes even more mainstream, it has the potential to become the script for many of the other sign languages as well.

Practically, I would like to start with this project on WMF servers using the wiki like software that already exists. I would look for programmers and or funding to enable MediaWiki to support SignWriting. I would look for funding to create a full implementation of the SignWriting glyphs in UNICODE.

Thanks,
GerardM

Thursday, April 12, 2007

Braille

When you cannot see, resources like Wikipedia are not available to you. When you cannot see, you need something like braille to make such resources available to you. You still have the issue who is going to convert text to braille.

RoboBraille is an organization that converts to several different formats that can be used by braille readers that work with computers. The software works by receiving an e-mail with the text that needs conversion.

I can imagine that this engine could work for a website as well.. Consider what a difference it would make when Wikipedia would be available in this way.. The converted pages can be cached like any other page. When there is an issue with doing this realtime or near realtime, it would be possible to do this for articles that are featured articles or that are part of a "final version".

I would love to see a solution like this to become part of the service that we provide.. We aim to bring all information to all people.. blind people qualify :)

Thanks,
GerardM

Tuesday, April 10, 2007

Linguists go Wikipedia

In a previous post I mentioned that as part of the funding drive for the Linguist list, the subscribers of the list were asked to vote with their wallet. They did, and an intern will be paid to organise an editorial update of the Wikipedia pages on linguistics. The idea is that the many thousands of linguist that are subscribed to the linguist list will ensure that notable linguists like Eve Clark and Tanya Reinhart will get a mention.

Consider, this is a leading list of some 14,649 linguists that will be urged to help us improve both the quality and the quantity of the coverage of the field of linguistics.

I am really exited about this.

Thanks,
GerardM

Saturday, April 07, 2007

A board of trustees or an executive board

In a post on his blog Tawker suggests that it was "widely reported why" Mr Wool left his job. This is not true. Mr Wool explicitly refrained from explaining why he left his job and in this way prevented a discussion both about the organisation and his role in it.

The organisational consequences in the statements made in this blog are also very much wrong. The Wikimedia Foundation is in the process of setting up an organisation. This effort is hindered by a lack of funding and an unlucky choice in staff. As the staff becomes more professional, the board will be able to distance itself from the day to day affairs.

Thanks,
GerardM

Friday, April 06, 2007

Vietnamese

In OmegaWiki you can have many of the labels used in the data part in your language as well. For this to function, you have to select a language in the "user preferences" and there have to be translations of the words involved. Words like Estonian, Georgian or noun, verb, adjective will be shown in the selected language.

As the user interface is one of the most critical aspects for getting buy-in, it is something I often spend time on. Given that I do not speak languages like tiếng Việt, it is not always easy to find the right translations. For this language I was not able to find Gujarati. More bewildering was that I could not initially find the word Korean, it was there as korean with "tiếng Triều tiên" as its translation.

All in all I have added more than 10 translations in this latest session. I am sure that people who speak this language will have an easier time doing the same job.

Thanks,
GerardM

Thursday, April 05, 2007

{{WOTD|petition}}

At OmegaWiki we have a word of the day. This word of the day is hopefully created before a new day starts. It is something that I find myself doing almost everyday. Most often I just pick a word at random. Today I picked the word petition.

Today I signed a petition, the Alan Johnston Petition, at the BBC News website. I think it is really sad when journalists are the victims of political violence. It prevents news coming from corners of the world. In this way you can prevent news going our or coming in. By stopping the flow of information, the risk increases that the "other" will be seen as the enemy.

I am afraid the Palestinians are shooting themselves in the foot. You can also sign the petition.. or think some positive thoughts about this.

Thanks,
Gerard

Tuesday, April 03, 2007

Linguist List's Wikipedia Update Vote

Linguist list is the premier mailing list that aims to provide a forum where academic linguists can discuss linguistic issues and exchange linguistic information. It is more than just a mailing list; it also provides fellowships to graduate students who serve in return as editors on the list.

Yesterday there was a message that following an initiative on the Russian Wikipedia, they would pay for a graduate assistant to work half time for one semester to coordinate the improvement of the field of linguistics in the English language Wikipedia.

When their community thinks it important, they have to spend $2000,- extra and this will pay for this initiative.

There is nothing that stops people not yet associated with the Linguist list to contribute to this funding drive as well. :)

Thanks,
GerardM

Thursday, March 29, 2007

Scripts on the Internet

I am back from the ICANN conference in Lisbon and, I have the T-shirt to prove it :) At a meeting the BSI proposed to create a standard that will describe how a mandated list with a ccTLD for every country or territory is to be produced in other scripts. This would be an industry standard that would eventually be adopted by ISO. Having such a list would be truly beneficial to get to the stage where the registrars for these registries know what codes to use. I have been told that a list has already been compiled for the Arab and the Cyrillic script.

On the BBC-news website there was a great story that explains why it is so important to allow for content in the language that people speak ..

One other thing I learned is that you can already have a .org domain name that is in an UTF-8 script. To me this is quite important because it proves that supporting UTF-8 is something that can already be done. It would make such a difference if the Internet was as functional as it is for us people who read and write the Latin script.

Thanks,
GerardM

Monday, March 26, 2007

Cryptography

Today I am at the ICANN conference in Lisbon. For me it is quite special to be here. One of the nice surprises was a gentleman editing on the Wikipedia article on Implicit Certificates. These are some special certificates for use in smaller devices. He mentioned that after writing this article, within the hour relevant edits were made to the topic.

The great thing for cryprographers is that Wikipedia provides both support for LaTeX and as relevantly, it can be updated when new developments happen. This makes Wikipedia an excellent environment to maintain information on this domain.

Thanks,
GerardM

Thursday, March 22, 2007

Citizendium and licensing

In yet another diatribe Dr Sanger informs us about the differences between Wikipedia and Citizendium.

The thing that struck me most are two things; contributors have to give a non-exclusive license to Citizendium AND the license is now to be the Creative Commons CC-by-nc license. The consequence is that only the Citizendium organisation can license commercial use, obviously for a price. They assume that as it is to be written by experts it will have value. It also means that once licensed, the licensee can do whatever.

When you compare Citizendium with Wikipedia, you have an English only versus a multi-lingual project. You have a project that covers almost everything and a project with a few thousand articles. You have a Free project and a project that is increasingly restrictive. You have a project that informs the world with NPOV information and a project that is to be written by experts.

Really I think Dr Sanger is doing a great job promoting Wikipedia by increasing the differences.

Thanks,
GerardM

Wednesday, March 21, 2007

Supporting the visually impaired

The content that is created in many Wikis is really relevant. Wikipedia for instance provides you with unsurpassed encyclopaedic type of information. This popularity can be deduced by Alexa's current ranking for today .. number 9. For us users and editors of Wikipedia it is like a roller-coaster; it is exhilarating :)

For people who are visually impaired, several of the Wikipedias have projects where people record the articles .. Really nice, really relevant.

Today I learned about new software promoted by UNESCO called Sakrament Libreader, it allows for text to speech conversion for English, Russian and Belarussian. Just consider, the English Wikipedia takes, according to Alexa, 53% of our traffic. There is a certain percentage people that are visually compared to the unimpaired.

Consider what would happen if the Wikimedia Foundation would support and promote this kind of functionality. Not only would many people be helped by this, it would also give an impetus to make this functionality available for other languages..

Thanks,
GerardM

Saturday, March 17, 2007

IATE became available

IATE or the Inter-Agency Terminology Exchange became available as a resource on the Internet. An introduction to IATE explains nicely what IATE is about; it is about the terminology of an organisation; the European Union. It aims to demystify the jargon used by the organisations of the EU.

In the past IATE was accessible as well; it proved popular and it was quickly hidden behind a password. This time you can get access by using the URL http://iate.europa.eu and it hopefully means that this resource is now officially available.

When you compare IATE to what it replaces, it is a massive step backwards from a copyright point of view. EURADICAUTOM was available under a much less restrictive license. It would be nice if the EU would steal a page out of the US book; most of the information provided by that government is available without restrictions.

Thanks,
GerardM


Tuesday, March 13, 2007

Google Summer of Code

Google has announced its third Google Summer of Code. This is an annual event where students develop on Open Source projects. This is definetly one of those activities that does a lot of good. It is one way whereby Google makes its mantra of "do no evil" work well.

For Open Progress, we have entered for a first time; we have a nice mix of MediaWiki and OmegaWiki based projects. All these projects are dear to us. We have shown Brion our list, and we are likely to work together on these.

What struck me is that when you apply for the GSOC, it is compulsory to have a mailing list. This is the traditional way of doing things. I am subscribed to many mailing lists. I think mailing lists suck big-time. There is so much repetition, the signal to noise ration is typically quite bad. I do not understand why people do not use a wiki to document and discuss.

I think this is one of those instances where software development proves to be conservative. When you follow the subjects you are interested in on a wiki, you can use watch lists to make a selection, you can use RSS to follow the changes on a low bandwidth wiki.

Because you refactor what is there, there is no need to repeat so much. When you have discussions that are getting out of hand, backrooms can be opened for those quarrelling. Maybe I am an idealist that I see it in this way .. oh well ..

Thanks,
GerardM

Sunday, March 11, 2007

Fon or Meraki .. I want the functionality of both !!

In the field of Internet connectivity, WIFI is the thing that keeps a road warrior and Internet junkie sane. It also keeps him poor. The amount of money some companies dare to charge for connectivity is tantamount to high-way robbery. Particularly at places where you have to spend lots of time, like hotels and airports are really unfriendly places. Particularly in airports it is galling; you have to be there so many hours in advance, you are often delayed .. it would make more sense to provide it for free and in that way keep the punters happy.

With many Internet organisations increasingly dictatorial in what you can and cannot do, with intellectual property organisations only interested in filling the pockets of the big companies, WIFI is a next place where the people can create a commons away from these established malpractices.

Fon is a WIFI sharing organisation where you provide access to other fonistas by sharing the Internet connection. In one of the more innovative approaches they are targeting the neighbours of Starbucks with free routers. The business model is that you can pay a small amount for access to the network.

Meraki is a WIFI sharing organisation where you provide Internet connection to an area using a mesh network. By including repeaters in strategic places, the area covered can be quite extended and, multiple Internet access points can be part of the same network. When you operate a network, you can determine who can access the network and even provide access for money.

I would like to have a mix of both. At home Meraki would be ideal because my ISP uses a different technology from the ISP of my neighbour. This means that I will improve the connectivity both for myself and for my neighbours. It also allows for providing truly local information. Fon would be ideal when I am away, with an increasing number of fonistas it means that me providing connectivity is the assumption of good faith that will find its sweet rewards.

Combining the two would be awesome. I would not hesitate long to go that way.

Thanks,
GerardM

Wednesday, March 07, 2007

Localisation of MediaWiki

There is a policy for new languages in the Wikimedia Foundation. One of the key things is that we want to prevent new abominations of projects where the language is not what is advertised. We have seen these in the past; one of the worst in this is what is called the Belarus Wikipedia; the people in control of this project prevent the use of the Belarus language as it is used in Belarus. This is a really awful situation and the language commission has asked the board repeatedly to act on this.

What we want to achieve is that new projects will promote cooperation in stead of establish division. Even though linguistically there is a substantial difference between the different forms of English, it is generally accepted that there will be only one English Wikipedia. This means that when the differences between two languages are less than for the different forms of English, the language committee is not likely to approve a new project.

When there is merit for a project in a new language, the people promoting this language have to show their commitment in the Incubator. This is where the environment in that language for the requested project is set up. What is expected, is that there will be a number of well written articles about different types of subjects, there should be a main page and, the most visible parts of the user interface should be localised.

The sad thing is that the aspiring projects cannot conform yet to the requirements; when a new Incubator project is set up, the message file for the new language is not created. What is needed is for one of the developers to create the necessary files so that the localisation can start.

For many languages in the past, there is no message files either; it means that the localisation is done locally and that this effort does not lead to the localisation of the MediaWiki software.
I am really pleased that for one language, Marathi, many of the messages have been imported into SVN by Nikerabbit. In a few days Marathi will be supported in all WMF projects. :)

I would welcome it when the Wikimedia Foundation gives the support of the minor projects a priority. At this moment it has none. The creation of message files in the Incubator for all languages and, when a language becomes a project, the inclusion of the first localisation into the MediaWiki software is the bare minimum. When we boast that we have localisations in some 250 languages, it should be a verifiable truth.

Thanks,
GerardM

About dictionary writing

Connel MacKenzie is one of the English language Wiktionarians who has had a big influence on the development of the Wiktionary project. He published his notions about what a dictionary should be. As I have posted a response to what Erin McKean, the current editor in chief for the NOAD, said in a presentation at Google, it is nice to write about in response to this as well.

For Connel, the project is about the English language. This is a big difference in approach to what Wiktionary is said to be about. He wants to limit it to those words that are part of the 600.000 most common terms. This is problematic because how do you judge something to be common and, what is common in one branch of the English language might not necessarily be common in another. He is of the opinion that "freak" terms should only be there in a sanitized form. To me it is important that a term is clearly and fully explained. When you "sanitize", it is not clear if the full meaning survives for someone who does not know the term. By disallowing multiple word entries, you loose the connection to those entries that are single entries in another language..

In his commentary, Connel writes about the restrictions that faces Wiktionary that are the consequence of its flat file format. You can not segregate different types of content when the basic technology does not support it. At that he would be better off being part of OmegaWiki as its technology allows for all the things he is looking for.

He hopes to get a useful Wiktionary when he has a dozen programmers available for such a project. At the same time he despairs because of "the current anarchy" it may take ten to twenty years..

I do admire the constructive work Connel has put into Wiktionary. I doubt that Wiktionary will ever become useful other than as a resource where you can look things up on the Internet.

Thanks,
GerardM

Tuesday, March 06, 2007

Upper ontologies

"An upper ontology attempts to create an ontology which describes very general concepts that are the same across all domains". As almost always there is a Wikipedia article about this. Given that the subject is difficult, there is since August 2006 a request to clean up this page .. I have to agree that the subject is interesting and potentially controversial.

For OmegaWiki, an upper ontology is one of those things.. An upper ontology defines the broad strokes, and while it descends downwards more and more aspects may be inherited from higher levels. This means that an upper ontology has practical implications. It also means that we will eventually have in effect an upper ontology by default.

As an upper ontology creates the concepts that are true across domains, Wikis for Professionals will want to define how their ontologies fit into the upper ontology. As they are bound to have overlaps with other domains, there will regularly be found to be in conflicts between domains. As one WfP needs to link into other domains, the question of primacy will raise its ugly head. One WfP was there first, but this other domain is not its competency... A new WfP does have the competency but is disrupts the existing WfP...

An initial selection of an upper ontology will be crucial, its evolution will be exceedingly important when OmegaWiki will is to be bound by the integration into it. Personally I expect that a different model will arise; one where on the one hand great care will be given to the evolution of the upper ontology while on the other hand functionality will be created irrespective of the upper ontology.

Consider, when you know that something is a plant, you can infer all kinds of things about it. It is not really necessary to know how it fits in the greater scheme of things. It would be nice, but it is not required. When a lot of functionality is determined in such a way, there will be the question how this will fit together; it is therefore my prediction that these two forces will eventually find a balance. As OmegaWiki matures, this balance will become increasingly stable.

Thanks,
GerardM