Wednesday, February 28, 2007

NIH Wiki Fair

The National Institutes of Health organised a Wiki fair today. There have been great presentations; the line-up of speakers and subjects was impressive, many of the presentations are on line and the whole presentation can be seen as a video as well ... all 5:41 hours off it...

For Wiki aficionados particularly interesting are a presentation of Dr Larry Sanger, and Dr Barend Mons. Larry is part of the Wikipedia history and Barend is part of the OmegaWiki presence.

Larry's presentation is self serving; he gives a selly presentation about Citizendium and he urges the people from the NIH to join in his project. That is in and off itself OK as he just what he has to do. As the raison d'être of Citizendium is in doing a better job than Wikipedia there are loads of comparisons with Wikipedia.. It is not surprising that you find no nicks when that is your policy, it can not be considered something that is "good" in comparison. Given that Larry finds it relevant to be known as Dr Sanger, I would expect a more scientific approach to his presentation. Given the small size of the project, it is not so strange to find that there is a lot of harmony around. This is very much the experience of all Wiki projects. It has been well documented in many scientific papers and, I am sure is Dr Sanger aware of this.

Being part of the history of Wikipedia and having created Citizendium as a reaction to what is considered "wrong" in Wikipedia means that Citizendium will always be compared to Wikipedia. I think that having started from scratch is daring.. It is the right thing to do. There are certainly more scientists working on Wikipedia, and given that Wikipedia has more than 1000 times more articles, it will be interesting to see what kind of attention this project will get.

Thanks,
GerardM

Saturday, February 24, 2007

Open Access

Given that I am very much in the open source and open content world, it will not be a surprise that I am very much in favour of open access and that I do not think highly of the patent system.

I think that it is easy to argue that patents kill. My argument goes like this; the pharmaceutical industry is documented to be only interested in patentable medicine. When a medicine is not patentable, they do not have an incentive in researching its application for particular purposes. Recently the university of Alberta found by accident that dichloroacetate (DCA), appears to suppress the growth of cancer cells without affecting normal cells. The problem is that there is no funding for doing the research to prove the efficacy of this medicine. It is therefore easy to understand that only substances that are patentable are considered in much of the bio-medical research.

I think that is is easy to argue that the classic business model for scientific publishing kills. The cost of reading scientific articles is staggering. The article published in Nature about what is demoed on wikiprofessional.info, for me it is a really exciting article, for me this really relevant article costs $30,- to read. I can not lawfully send it to my mother to read. I can only tell her. Scientific journals assume that the only people who need to read their papers are scientists; the scientific libraries will have a subscription and that is how it has always been. Except that it is not true. Laymen have as much a need to read bio-medical science papers; they or their loved ones are affected by all kinds of afflictions. They have a need to understand what is happening to them. Often doctors do not know all the details that are available and , there are enough documented cases where research on the literature by laymen made a difference. One famous example is known as "Lorenzo's oil", a story about a disease called adrenoleukodystrophy or ALD.

Not only laymen are denied access to literature, many universities cannot afford the cost of literature. With the science results owned by business, it is important to understand how vital the whole "open" movement already is. A new, a different business model is needed and many organisations are developing these. The big challenge is for science to become science again; to be able to research without having the public pay again and again and some companies walking away with only an eye for profit.

The challenge will be to find a balance that does justice to all parties involved.

Thanks,
GerardM

Thursday, February 22, 2007

Mother language day

Yesterday, it was the International Mother Language Day. There is a program in Paris with all kinds of speeches and happenings. I read the program, I know some and I know about some of the people involved. I wish I was there. I wish they recorded some of the speeches, presentations.

One thing I find funny is that UNESCO indicates that there are some 6000 languages where SIL indicates the existence of over 7000 languages and where the ISO-639-6 will know at least 25.000 linguistic entities .. :) I wonder if they need to be emancipated so that UNESCO will feel a need to at least acknowledge their existence..

It would be cool if International Mother Language day had hit the Internet; a podcast would have been nice ..

Thanks,
GerardM

Sunday, February 18, 2007

A root canal treatment anyone ?

Yesterday, I had an appointment with a dentists specialised in the noble art of endodontics. I had a molar that had already had a root canal treatment before and it needed some more work. I was nervous. Actually I am terrified of dentists and consequently I am my own worst enemy as a person suffers most from the suffering that he fears and that often never materialises.

So I went to my appointment and I was wearing a Wikimania 2006 t-shirt. A person waiting in the reception reacted to this and, we got to talk about science, creating educational content in a collaborative way, licenses and copyright, OmegaWiki, the Nature article and the Wikiprofessional demo.

When it was my turn to sit in the "chair", the doctor had heard much of the conversation and asked several questions.. There was no time gap between being operated on and talking about the things that are so dear to me. One big difference between what an endodontist does and what a dentists does is in the tools of the trade; an endodontists uses microscopes and a lot of digital imagery. It really gave me the feeling that I was operated on. Anyway, after the operation I gave a demonstration to my endodontist.

The good news is; I did not have time to be nervous.. Two more people may have a good look at OmegaWiki. It feels good even though my molar is still sensitive .. :)

Thanks,
GerardM

Friday, February 16, 2007

A friend of my is a scientist...

A friend of my is a scientist. He is a terminologist. He has published a lot of papers and, he is considered one of the best in the field by some other people I know.

I told him about wikiprofessional, what kind of things we are doing in the bio-medical field. What we could do for other fields as well when we have the terminology available to us. This was two days ago. Today I learned that much of his work was once available on diskettes. These diskettes were no longer there .....

I discussed this with another friend. She has some great OCR-software. We know each other, so when the terminologist scans his paper paper, the translator can translate it from a analog into a digital format. The terminologist can define his terminology, his papers will be known but as relevant, we will have started supporting the terminology of terminology in OmegaWiki.

It is great to have friends ...

Thanks,
GerardM

Friday, February 09, 2007

Cleaning a bit on the sk.wiktionary

Many of the Wiktionary projects have a moribund existence. At some stage people worked on it. They left the project, things changed and where never properly taken care off. Today I was in an extended chat and I was not the lead talker, so I had some time to do some stuff that does not take much attention.

I cleaned up a bit on the sk.wiktionary. At some stage all Wiktionaries supported proper casing. This meant for the acronyms like ADHD that they became aDHD. Somebody needed to go and fix these so that they became properly ADHD. Today I fixed more than sixty acronyms that started with a, there are many more of these that need fixing and, not only but also on the Slovak Wiktionary.

I am really pleased that at OmegaWiki we only need to fix things once..

Thanks,
GerardM

Thursday, February 08, 2007

The Shtooka recorder

I was told to have a look at the Shtooka recorder. It is a tool that makes it really easy to record the pronunciation of words. You provide it with a list of words, you provide the meta data and, it creates an .ogg file for you in the directory of your choice. It makes it really easy..

For me the challenge was to understand how to use it. There is no user manual and mostly it is self evident. So when I understood that I had to press on the "record" button, it started to work.

This software does do wonders when recording words that are provided in a Latin script. I tried it to record a Persian word, کم حرف and it failed. I checked it with something Russian and it failed as well.

For me it is a big improvement, I did some 15.000 words with Audacity.. It would have saved me a lot of time when I would have had Shtooka.. The next thing to automate; the uploading to commons...

Thanks,
GerardM

Friday, February 02, 2007

Ten things you want to know about dictionaries

I met Erin McKean at the Wikimania 2006, I loved her presentation then and I was really happy when I found her presentation in the Google Author series titled "Ten things you want to know about dictionaries". I loved it, I have seen it twice now. I may even have added value to it by adding the "dictionary" label to the presentation. Here I am going to react it with my OmegaWiki hat on. So yes, please watch the presentation (almost an hour and well worth it) as I hope it will improve the understanding of my reaction.

There is no one dictionary; it is a tool
OmegaWiki as a resource is very much a child of the Internet; consequently it has the potential to configure its use. When people do not care for particular information; they should be able to make it invisible or turn it off. In a way it is like the cordless drill, by replacing the drill with a different thingie it becomes an other tool. The same is true for pronunciation; we love to include IPA, but we can also record pronunciations this way people do not need the understanding required when reading IPA.

Please read the "front matter"
People indeed assume that they understand tools like dictionaries and wikis for that matter not to RTFM. For a consumer good like a read only lexical resource, it is pretty safe when the introductions have not been read. As OmegaWiki allows people to add/edit to the information that is in there this proves to be much more problematic.

Inclusion in the dictionary is because it is useful
As we do not have all the functionality that we need to be a credible lexical resource, this is very much a state we hope to get at. However, our aim to include all the lexicological, terminological and ontological sounds pretty like megalomania. Our standard excuse is that we use is that is already less problematic because we only want to do this once and this is where we came from. The data will only be useful when there are people who care about particular categories of data. I am totally with Erin that only data that is useful should be included. Getting rid of unnecessary cruft is hard work.

Horrible words make it in their too
I am a fan of swear words in dictionaries, particularly when there is some etymology to it. Most often people use swear words as an expletive without much understanding for their actual original meaning. As English is for me a second language it is relevant for me to understand why I would rather be a bigot than a racist or someone who discriminates.

The other part of horrible words are those words that are actually used and offend the aesthetic sensitivities. In several medical resources you find stuff like MALARIA and Malaria. UGLY. However as it is useful to these folks, it makes sense to include them anyway. As they are exactly the same as the preferred English expression of malaria, it does not hurt.

Words like "irregardless" well being a non native to the language, I just want to be able to find them.

You have to look at all definitions to find the REAL meaning
The way the New Oxford American Dictionary does this is exquisite. They use a core sense / sub sense approach. To me this seems an approach that is very much language specific. For OmegaWiki to have such an approach, it will need quite a lot of thinking on how to build this.

Approaching the understanding of an expression with core senses / sub senses could be one way of stretching the number of concepts that people can juggle with. For Operational Definitions there is currently this practical limit of some 7 different meanings.

Dictionaries have a sell by date
For OmegaWiki this is not an issue as it is web based resource. However, the same issue still applies; a Dutch book printed in 2003 will use the orthography of 1995 and not the 2005 orthography. People will still want to be able to understand what this word means; annotating them as not being the official spelling since 2005 is relevant.

When words are tagged as used in earnest up to a certain date, I could even include 15th century German words and not have people be confused.. filtering would also help here ..

Facts are good
Referring to actual usage seems obvious, what we intend to do is link OmegaWiki's content to Wikipedia, this will be the most obvious resource to start of with; we aim to have a Wikipedia in all languages. The language is modern usage, so given our Wiki credentials it is the obvious corpus. I totally agree when it is said that we are limited in this way; it is however a great start. When we gain a community with people, organisations that introduce us to other resources, it will be great.

At this moment OmegaWiki is still very much like a "stamp collection"; it is a nice collection, and at some stage it will even become useful.

What we do is like an iceberg
The work done on the New Oxford American Dictionary may be in preparation for the moment when other information like thesaurus information will be included. For OmegaWiki, including information from thesauri is what we did from the start by including the GEMET data, OmegaWiki at this moment is very much "you get what you see", there is little of an iceberg yet. In a way this prevents the usefulness of our data because there is often too much to take in.

Etymology
Technically etymology is one of the hardest nuts to crack. When a word has its root in Latin, it often came to the English language through a French or Spanish connection. I wonder how the NOAD does this, indeed I do not have a copy so I have not read the "front matter" either :) .

Neologisms
At this moment there is not much of a problem about neologisms yet. Our community is still small and sane.. sort off (you must be weird to involve yourself in a project like this). My current thinking is that this problem can be solved using annotation. When a word is tagged as "Neologism; this word is not used except by the author" it will be pretty devastating to the prestige of the word and or the author on OmegaWiki.

The bonus: Using Google and other resources
Erin explains that the resources for building a resource like the NOAD are hardly as much as she would want. Given that this is true for a successful resource for the American English market, consider what this means for languages like Kituba, Stellingwerfs or Seeltersk. Consider what it means when you want to use a translation dictionary for such languages.. These resources become a reality when there is the necessary cooperation; this is what OmegaWiki hopes to achieve.

There is a need for people to work on their terminology and indeed we would like to include the terminology of falconry, tennis, and ships. It will happen when it does.

Q&A
*OmegaWiki does want to include out of copyright content as well. For us it is a start. Collaborating with for instance WordNet would be however more important and relevant.
*Proper names; yes we want them; we have George W. Bush already for quite some time.
*We would retire words by indicating them with a date indicating when they went out of use
*Circular definitions are even more problematic in OmegaWiki, this is a great example why.
*Context .. yes, I wish this was a problem that we have to deal with.. we need more functionality
*Print versions .. this is at this stage no issue. Nobody has indicated that they want to work on this.

Conclusion:
I really enjoyed Erin's presentation. It helps me to get my mind around issues that have not popped up for OmegaWiki. As many are quite will make their appearance, it is best to be forewarned, it allows us to get forearmed.

Thanks,
GerardM

Thursday, February 01, 2007

Google defuse the googlebomb .. GREAT

When you are of the opinion that George W. Bush or "Shrubya" is a "miserable failure", you could find confirmation for this by googling for this and this truth would be on top both in the Google, Yahoo or Microsoft search engine. Effectively it is a prank. It is not what you really want to find and Google announced it has worked on an algorithm that will prevent a Googlebomb in future. Effectively making the word a misnomer it is now more correct to call it a Yahoobomb or even better, a Microsoftbomb.

In an article in the Guardian, the fear is expressed that by manipulating rankings in this way, Google will exert its power and be able to manipulate what is seen as true. At issue is that the reason why Google defused this bomb was because people believed it to be true because Google said so... (there is no such thing as common sense as common sense ain't common)

In many content projects, bots create links to websites they hope to make more relevant in the eyes of the search engines. This type of vandalism resulted in a backlash where Wikipedia now indicates to search engines to disregard any and all links and thereby invalidating the basis on which search engines operate. When Google were to have algorithms that filter and punish this type of SPAM, it would lead to a more sane environment.

Google is open about its intentions. Microsoft is open about its intentions as well; as long as it discriminates against its competitors like Wikipedia I will be happy to use Google knowing that Google is kept honest by having competitors.

Thanks,
GerardM

Wednesday, January 31, 2007

To skype or not using Skype

Skype is one of the most important tools for me. I have loads of people I do contact this way, they are all over the world. OmegaWiki is possible because of using VOIP is easy. I talk with people on every continent except for Antarctica. For me it proved to be an essential tool.

The problem is that several of my most relevant contacts do use Windows and the quality of the Linux implementation does not cut it. Skype freezes the sound system erratically or it just does not install. One of my friends wants me to use SIP; he threatened me not to talk to me any more.. Remember, this is a relevant contact to me.

We had a look around and we found that we can connect a SIP service to Google talk. This worked great; I can connect using Google talk. If there is one thing lacking, it is that I can not initiate the connection to the SIP servers, they have to connect to me.

As I needed to find out if SIP would work for me, I installed Gizmo project. It basically does everything that Skype does, it uses a Standard protocol, in stead of a proprietary protocol.

My conclusion; to skype I do not need Skype. I would probably be better off when all of my contacts would not be using Skype at all. The message to Skype is simple; the one thing they have is a user base. The support for Linux sucks and this makes people look for pastures green. So, FIX IT !!!

Thanks,
GerardM

Monday, January 29, 2007

Scientific publishing

Scientific publishing, I read a year ago, is the most profitable part of the publishing business. It has increased its profit continuously more than inflation for the last fifty years. To be beneficial scientific publications need to be an essential tool for scientific development but they have degenerated to a point where many University libraries cannot afford the price for what should be an essential tool.

Newton said that "he could see further by standing on the shoulders of giants". In this day and age where it is effectively not possible to read all the literature published within a scientific speciality there is a crisis. The crisis is that it is impossible to read everything anyway and much of what is out there can not be read because it is impossible to afford it.

When a system is broken, inertia will prevent things from changing. When a tipping point is reached, things break, they break fast and if there is nothing to replace it, things break badly. For the publishers the prospects are bleak. They have made it plain that they plan to play dirty; they have hired a PR guy that is known to defend the indefensible. This can only bring the moment when things break that much closer.

By going for half truths and lies, they will alienate the very people they rely on. They will destroy this carefully created façade of an industry that is a benefit to science. They will destroy much of the goodwill that they depended on. In a way it is sad, in another way some call it evolution.

Thanks,
GerardM

Saturday, January 27, 2007

Dictionary content in Wikipedia

In those days, there were articles in Wikipedia that were short. They were deemed to be too short, they were only a definition.. they were not even a stub. They got deleted. People started to add the kind of information you find in dictionaries. They got deleted. Wikipedia is an encyclopaedia was the argument. The quarrels were intense and at some stage Brion Vibber created Wiktionary, a project aimed to include all words of all languages.

Today I was back in the "kroeg", the Dutch Wikipedia village pump. The same argument still exists; there are still articles that are little more than definitions. They still get deleted. These deletions are protested against. In a way it was comfortable, I was back and it was as if nothing had changed.

The same old argument is used; copy the information to Wiktionary. What these Wikipedians do not know or forget is that Wiktionary has its own format and stuff that does not comply with the rules gets .. deleted.

Thanks,
GerardM

Friday, January 26, 2007

Sample sentences

One of the functions in OmegaWiki that is currently not popular is the ability to add sample sentences. Sample sentences are a type of annotation on the level of synonyms and translations, they are "text attribute values".

You may appreciate that I look at many resources to learn what functionality might be relevant. What I am starting to appreciate is that there are two relevant approaches to sample sentences. One way to approach them is to have a modern sentence that demonstrates the word well. The other is to have quotes from great authors like Shakespeare. The "bard" may be a great writer, but many of his sentences are hard to appreciate by students that are new to the language. Often there is also some semantic drift that makes it even harder to understand what is meant.

For linguists, sample sentences of the earliest use and the later use of a word are important. There are therefore two constituencies for sample sentences. The question is how to deal with these; having annotations with the quotes make sense for the linguistic inclined ones. There may even be a need for a flag of "modern usage".

One approach that I really admire is the one taken by Logos in their Logosdictionary; they have functionality that shows you the word in a corpus. This way the sentence is never too short as it is sustained by the corpus itself. It would be cool to emulate this with Wikipedia content.

Thanks,
Gerard

Wednesday, January 24, 2007

LinkedIn

LinkedIn is one of those social networking websites; it allows you to build a network of people that you are associated with in a business sense. One of the persons I am associated with is Polish and his name contains the ń.

This should be no problem for a networking website; what software does not support UNICODE ? The name of my friend is spelled incorrectly and, for my network of people, it is essential that people can life every where on this globe and have their name in whatever character set.

I have asked LinkedIn to look into this. For me it is important.

Thanks,
GerardM

Tuesday, January 23, 2007

Google ads

Google has noted that I have something with Tagalog. Their adsense targets me with an advert for translators who offer there services for Tagalog. I find it funny and note worthy as we do support this language for a month now in OmegaWiki. I am probably one of the few that have added Tagalog translations to OmegaWiki. It is uncanny because this is the only thing I am aware off having done with respect to this language.

When you understand this as a request for more people helping out with Tagalog, you are also right :)

Thanks,
GerardM

Saturday, January 20, 2007

Toilet and doorknob

Terminology is everywhere. The need for translations is needed all the time. A dear friend of mine, is an engineer and is improving her Spanish. She does this with regularly speaking, chatting to someone who does speak .. Spanish and wants to improve his .. English. Both of them are into wikis and, I learned that some technical material in English was send and this resulted in this beautiful article on how to fix a running toilet. OmegaWiki does have the word for toilet but do we have flush, tank, lid, flapper, bowl, float, valve, overflow tube ??

The wikiHOW article is in need of translation for instance into Spanish, because there are running toilets wherever you find toilets. OmegaWiki would love to have the terminology... it is a doorknob (another word we should have :) ) that will open many more opportunities for adding terminology to OmegaWiki.

Thanks,
GerardM

Wednesday, January 17, 2007

Standards .. we need them

I had another conversation with someone who has a big interest in the Apertium machine translation software. OmegaWiki has something to offer to offer to tools like Apertium; it is a database, it intends to have all words of all languages and last but not least it's data will be available under a Free license.

To be of relevance to tools like Apertium we need to have conjugations and inflections. This is something that we have always planned to include. The new thing for me is that it is desired to have morphological information as well. With the current costs of computer hardware, including the data is not that much of an issue; as it is only text this will not really amount to a big increase in hardware costs.

I then asked the question; is there a standard format for morphological data.. Apparently there is no standard for it. So what to do when there is no standard. You can select one way of showing morphological data amd this will piss people off or you can create a new standard and piss off every one.

The good news is that we first have to deal with the things that come first; being able to include conjugations and inflections. So there is some time before it actually becomes relevant. The thing is it will, and it is likely that we will do something. When we select one way, I hope it will be flexible enough to allow for conversion to other formats as well.

Thanks,
GerardM

Tuesday, January 16, 2007

Blogger out of Beta

The Blogger software is no longer a beta product. I have been blogging for some time now with Blogger and now I can use this new functionality.

With the software going out of beta, a lot of great new functionality has become available to me. I am pleased to the extend that I want to say, "thank you". I like it that I can label my posts, that the options can be set for each individual post. That users can select the posts that have the same label. The way it is possible to customise the layout of my blog it is so much more WYSWYG. :)

Really, thanks,
GerardM

Sunday, January 14, 2007

Some homegrown statistics

OmegaWiki grows, the statistics provided by Kipcool show that quite clearly. Both the number of DefinedMeanings and the number of Expressions show a similar rate of growth. The current ratio between them is 14,49 Expressions for every DefinedMeaning. At the moment it is slowly decreasing.

This is in a way counter intuitive; the number of languages is steadily growing and often more translations are added to existing DefinedMeanings. It demonstrates that many concepts are added without too many translations. These translations are likely to be added at some stage.

Particularly the DM curve I expect to flatten soon. I expect that realistically there is a limit to the number of concepts out there. With the bio-medical data that we want to include, I expect that the ration will get a beating but this will slowly but surely creep again. After all the ISO-639-6 expects to include at least some 25.000 linguistic entities ...

Thanks,
GerardM

Saturday, January 13, 2007

Two friends changing their operating system

Two friend of mine changed the operating system on their laptop this week.

One changed from the top of the Windows Vista with all trimmings to Windows XP. He was not willing to pay for new applications. Vista is incompatible with previous Windows software I was told. This is very much a non event, it is just interesting

The other changed from XP to Gentoo Linux. To me this is a pain. He is not able to get skype to work. To me this is a major pain. What makes it worse is that as I consequence I cannot conveniently call him any more. His reaction is that I should not use proprietary software and use SIP compliant software myself. I already have two clients; Skype and Google talk and I am sad that both of these do not support Skype.

I got into a bit of a fight about the subject. To me, there is no such thing as a Linux. There are hundreds of distributions and all of these have there own requirements. I hope that that the LSB
will provide the necessary integration so that applications will work never mind the distribution. At this moment Linux is still very much an operating system for developers. Where it is a damnation when I utter this word, it is high praise for my friend..

Really, to me a computer is first there to be used. It has to be easy..

At this moment I would like to have a SIP client under Windows.. The thing I would like best is for Google Talk to finally support SIP..

Thanks,
GerardM

Wednesday, January 10, 2007

Elder Futhark

Elder Futhark is the oldest form of the runic alphabet, used by Germanic tribes for Proto-Norse and other Migration period Germanic dialects of the 2nd to 8th centuries. It has some 24 runes and these and others are included in the Unicode.

I can easily imagine that people will want to include words in Proto-Norse in OmegaWiki. There is probably not much actual knowledge about this language. There is no ISO-639 code for it.. Well, when we really want to include it, we will. It will look funny on may people's computer; who has the fonts for runes ??

Thanks,
     GerardM

Sunday, January 07, 2007

A toe, a finger or a "prst"

In OmegaWiki, we have a problem in that we are busy beefing up the content. And sure, I am to blame for it myself as much as anyone else. The problem is that in order to achieve a substantial size in translations and thereby achieve some relevancy, it is easy to forget about the basic underpinning of OmegaWiki. The problem is that in several languages, there are no words as precise as the words in others. In Croatian, and some other languages, prst is the word used for both the finger and the toe. The consequence is that there should be another DefinedMeaning for prst describing it as "One of the five extremities that can be found on a hand or a foot".

Creating such a DM is not a problem however, I can not define a DM in Croatian. I can define it in English, Dutch. The issue is that a DM is defined as a combination of an Expression and a Definition in the same language. It does not stop me from creating a DM.. it does mean that we need to consider this scenario as well.

Thanks,
GerardM

Friday, January 05, 2007

The OmegaWiki mantra

According to Guy Kawasaki every start up organisation should have a mantra. It should be short and to the point. Sabine ownes the perfect domain and that domain is the perfect mantra for OmegaWiki; "Words and more".

I think it is the best way of describing what we aim to do with OmegaWiki. I am sure that Sabine agrees..

Thanks,
GerardM

Sunday, December 31, 2006

Collaboration

It is the last day of the year and, it is a great moment to consider what to do in the New Year. For me the one thing that will typically make the difference is collaboration. What we aim to do with OmegaWiki will be a success when we aim to be inclusive. This will allow everyone to benefit from the mutual effort on the same data.

The economies of scale really will work to our advantage when the work on the data is shared. When people and organisations work together, tipping points will be reached that will enable the application of the data that would otherwise not be really possible.

Today I learned that it may be possible to collaborate with Logos. This would be really great; Logosdictionary is already relevant, it has a community that make the daily logosquote a rich reality. It has a children's dictionary, a conjugator and more. Truly the prospect of such a collaboration is significant, the challenge will be to make it happen.

2007 will be an interesting year .. :)

Thanks and "gelukkig nieuwjaar!"
GerardM

Saturday, December 30, 2006

Statistics

The Marathi Wiktionary has 121 articles. OmegaWiki has 120 expressions in Marathi. I had a look at the recent changes of the Marathi Wiktionary, it was full of changes to the MediaWiki messages. Consider, the same work was probably already done on the Marathi Wikipedia. When these updates are done in this way, people will not be able to benefit on other projects that are of interest to people that read and write Marathi.

It is really relevant that the work done on languages like Marathi count. BetaWiki was a place that functioned well, however as it was not part of the MediaWiki projects, it was ignored by some of the developers. With the inclusion of the software written for BetaWiki in the Incubator, there will be a more obvious place for the localisation of MediaWiki. It will be a place that is easy to understand for translators. The work done there will benefit all the projects where people are able to use their language for their User interface.

Thanks,
GerardM

Wednesday, December 27, 2006

Playing a game

It is the festive season and families come together. Adults talk and kids play. The kids in my family play games, computer games. I was asked if they could use my computer to play a multi user game. I did not mind, I made them promise me that they would uninstall this game when they are done / before I am to leave.

The question; is this the type of thing that Industry considers illegal.. It must be because no money changed hands. If this is indeed illegal, where is the fun?

Thanks,
GerardM

Lies, damned lies and statistics

There was an e-mail on the Wiktionary mailing list where a comparison was made between the English language Wiktionary and OmegaWiki. It was based on an analysis by Zdenek Broz. It compares the number of translations between the two projects.

The numbers may be right, but there is so much more to the English Wiktionary that it should not be reduced to such a numerical comparison. There is the Wikisaurus, there is a lot of etymological information, frequency lists, rhymes and most relevantly there are a lot of people who make it a great project.

It will take a lot more work before OmegaWiki can include all the information that is in en.wiktionary and it will take even more work before this information will actually be in there. I think this may happen but we do not need a competition for that. Both projects have there stong points and when we are able and willing to learn from each other it will be awesome.

Thanks,
GerardM

Friday, December 22, 2006

A Christmas story about traditions

I really enjoyed this wonderful Christmas story. It is about centuries old traditions and about personal traditions and how they change.. I really enjoyed and want to share it with you ..
Thanks,
GerardM

Wednesday, December 20, 2006

Running in front of the pack

MediaWiki is the software used by the projects of the Wikimedia Foundation. If there is one thing MediaWiki can boast about, it is the amount of localising that has been done. There are currently projects in 250 languages and, many have been localised somehow somewhere.

The Wikimedia Foundation has a big need for money. If ever there was an organisation that took care of the money it received it is the WMF. Wikipedia has according to Alexa the 12th traffic rank at the moment and the growth of their projects is something like doubling every four to six months.

The WMF has an ongoing funding drive, they created software to manage this in Drupal. The functionality is good for some languages but for others Drupal has not been localised. Some will say: it does a good job for my language but for most languages Drupal is just not up to the task. MediaWiki is far ahead of the pack when it comes to localisation and consequently there are no tools that will support the languages that are needed by the Wikimedia Foundation.

When Drupal gets more localisation done, it will find that in the same time frame the WMF will have added more languages. There is a moral question here as well. Is it acceptable to use tools that convey an important message that do not support the languages that you need..

Thanks,
GerardM

Tuesday, December 19, 2006

Glyphs, fonts and the need to support them

In OmegaWiki (formerly known as WiktionaryZ) we want to support all words of all languages. Some languages have scripts that are not supported by the Operating System that you use. This can be bad or it can be really bad. It is bad when you have to find and install a font. Where to get it, what does it cost. It is really bad when there is no font. It can also be extremely bad when there the glyphs have not been defined in UNICODE.

In the Wikimedia Foundation there are projects that do require another font. Khmer, Laotian come to mind. For the Ripuarian language the situation is extremely bad; there is an official orthography that defined characters that do not yet exist in UNICODE.

Today I learned about a really nice project called Dejavu. It is an open source project that works on the creation of fonts. They do good work but there is a long way to go before they will have tackled the Chinese and Indian languages.

For OmegaWiki projects like Dejavu are important because they will ease the use of those people who are interested in seeing all the information that it will contain. UNICODE and the work done by Michael Everson are as important because without the glyphs being defined they will not go into fonts and without fonts we can not have all words of all languages.

Thanks,
GerardM

Monday, December 11, 2006

A tribute to the creator of a language

Many people are not impressed when non natural languages are mentioned.. WiktionaryZ is by its very nature more inclusive. That is to say all artificial languages with sufficient recognition will be welcome. One class of languages that WiktionaryZ will not include are programming languages.

Because of my background, I am personally very much interested in these languages too. Today, I learned from an article on the BBC-website that Grace Hopper was born 100 years ago. Rear-admiral Hopper was one of the most influential people in the development of computing. She is certainly one of my heroes.

Thanks,
GerardM

Wednesday, December 06, 2006

Preparing for a conference

In two weeks on December 14 and 15 I will speak at a conference in Vienna. I was surprised when I was asked to speak at the Language Standards for Global Business. What is it that I can bring to the table. Yes, I proposed a Wiki for Standards at the Berlin conference and yes I am a member of the Wikimedia Foundation language sub committee and yes WiktionaryZ is a big user of standards. I am still shocked and awed that I was asked.

Now that I am getting used to the idea I find that standards were increasingly taking up my time. Standards are crucial for the projects I am involved in. When they fit a need, you do not need to explain and argue why individual choices were made, you refer to the standard. When you export data and you implement a Standard for the format like TBX, LMF or OWL, you hide the complexities of the database and you provide the export in a stable, mature way that allows people to build upon.

The problem with standards however is, that there are so many of them. Also many of the standards work cross purposes by focusing on single issues. This lead to separate standards for the Internet, for libraries all indicating languages.. This plethora of standards and requirements prevents interoperability. It also prevents the general adoption of these Standards.

As I learn more about Standards, I find myself with WiktionaryZ at the cutting edge. How to publish content for a language like Bangubangu ? As far as I know there is little or no content on the Internet at all. I am pleased now to have the Babel templates for Bangubangu. But as the content grows, how do we get the search engines to find it, it is here where appropriate Standards can make a difference.

In preparation for my presentation I have looked at other conferences and I find that many are very much driven by commercial needs. Needs that do not necessarily take into account what is in the long tail of the industry that is represented. The maturity that can be found in translations between the languages of economic power houses like America, Japan and Germany are worlds apart of the African court rooms where the defendant is lucky when he understands the judge or a witness. Here there is often no translation and the tools that are available in the translation industry are not available even for the translation of court papers.

My problem for the conference is that there is so much that I would like to address that I have to concentrate and make what I will say count. The good news is that http://wikiforstandards.org is open for business. When people interested in standards take an interest and collaborate we may address all issues and get a better understanding what all these Standards are there for and more importantly, how they interrelate.

The challenge will be to build a community that understands that it is only collaboration that will make their Standards relevant and integrated with other Standards that are not relevant to their business

Thanks,
GerardM

Tuesday, December 05, 2006

Another certificate but no tulip

In my living room I have a vase with tulips. Every one of these fabric tulips stands for a certificate that I earned in my computer career. To me it expresses that in order to remain up to date you have to work.

Last week I was invited to a lecture by Marshall B. Rosenberg in Rotterdam. Mr Rosenberg is credited with developing a method of communication called " non violent communication". I attended, and I got a certificate to prove it. It does not rate a tulip because there is no achievement for me yet.

In the hand out we got, there were list of feelings and needs. It was mentioned that the English language was not made to express feelings and needs, the Spanish language was mentioned as being much richer when you want to express either. When you think of it, it is not surprising when language can be intimidating. The choice and the appreciation of words makes all the difference. Like it was explained during the lecture, in order to change your language you have to be aware about what you say and how the effect is on the party that is on the receiving end of that language.

In this discussion about male domination that I mentioned in my previous post, the language used is a major contributing factor to the unease that is being felt. When people do not perceive that it is their very words that makes others feel uneasy, angry even discriminated it is very hard to come to an understanding. Yes, it may be that for some the English language is a second or even tertiary language, it does not negate the effects these words have; it is at best an explanation not an excuse. Bullets do kill, words do hurt.

Thanks,
GerardM

Wikichix

On the mailing list for the English Wikipedia there were some women who articulated that a sizeable group of women feel not comfortable in the en.wikipedia community of editors. To alleviate this issue, they decided to create a community of women called Wikichix.

Angela, who is one of the best around for creating communities, set up some infrastructure including a mailing list for those interested in joining. A lively discussion started about this. Many men denied that there is a need, and that it is appropriate to have such a self help group. Several of them expressed that they feel excluded and even discriminated against. Several women including Anthere, provided graphic evidence of how women are dealt with by some of the males that think nothing of making disparaging remarks qualifying them as "jokes".

Even though it was said time and time again that this was to engage more women to become part of the Wikipedia mainstream, some people could not accept that what is good for some does not need to be an affront to others. Sadly, Anthere has now asked for the mailing list not to make use of WMF infrastructure. :(

Personally I feel it as a loss. A loss because it does not help to engage more people. A loss because it may even prevent the engagement of people who are not part of the Wikipedia community and all this because of the boorish behaviour by some.

Thanks,
GerardM

Tuesday, November 28, 2006

A new language on the Internet ?

I received a mail that was forwarded to me by Martin Benjamin (the Kamusi project). The mail was by a gentleman who wants to promote his mother tongue, the Bangubangu language. This is the first time that written material in this language has been created. The first project is a book in English, French and Swahili called “Teach yourself the Kibangubangou”.

I am thrilled with initiatives like this. The book is there now what to do next. Do you print it with an organisation like Lulu? Do you advise to make it a Wikibook? To what extend are concepts like copyright and licenses relevant and understood ..

It would be great to make this project succeed and have kids learn their mother tongue. There are stumbling blocks. Google does support less languages then the Wikimedia Foundation; Google does some hundred and The WMF some two hundred and fifty. One of the problems is that there is not much material in any but the bigger languages and Google does good by already doing this much.

To make an impact, I think it is crucial for material in particularly the endangered languages, to be tagged correctly. This gives Google and the other search engines a fighting chance to function on little material. For the Bangubangu language, there is no proper tag yet. The question is if the IETF will consider to do something about it. They are religious in their belief that the ISO-639-3 is not approved yet and that the bnx code therefore is not to be used. At issue is that there is a need that presents itself for this language.

Thanks,
GerardM

Sunday, November 26, 2006

Kanab Ambersnail

According to a heading on the English Wikipedia mainpage, the article about the Kanab Ambersnail was the 1.500.000th article. I think this is very spectacular. Personally I find it funny that I can write about it before I have read about it on slashdot :)

Congratulations to all who make Wikipedia so special.

Thanks,
GerardM

The I&I conference

Last Wednesday and Thursday I was at the annual I&I conference in Lunteren. This conference brings together many of the ICT (information & computer technology) coordinators of Dutch and also some Flemish schools. I have had the privilege to be for the second time. This year we had beautiful new Wikipedia folders, Marc Bergsma demonstrated a OLPC motherboard; it got a lot of interest particular when it was learned that a full system can be available for Dutch schools for the school year 2008/2009.

One great project I learned about is called TEEM. TEEM is a project where teachers evaluate educational websites and CD/DVD based resources for their use in classrooms. The way it is organised is such that I can understand why British teachers trust it as a resource. The way that it is funded however is one that creates a sad systemic bias. Reviews are paid for by publishers the consequence is that open content course ware will not be evaluated. This funding model also keeps out those publishers that do not pay for a review. This governmental political choice was to encourage innovation and competition. It is not unreasonable to suggest that it effectively costs the British schools more money as they do not learn about what is available for free.

Thanks,
GerardM

Friday, November 24, 2006

Whose language is it anyway ?

Much of Microsoft's software has been translated in the Mapuzugun or is it the Mapundugun language. They did this in consultation with the Chilean government.

The Mapuche people have gone to court because they disagree that Microsoft or the Chilean government had the right to do this. The language it is claimed is theirs, and the translation was done without consulting the Mapuche people.

Many people do not understand or know what is behind this. Why would people be opposed against becoming part of the digital world? To me, the key thing to appreciate in the reporting is the accusation of "violating their cultural and collective heritage". There are competing orthographies for this language and one is strongly favoured by the government.

So if I remember things well, this has everything to do about the involvement of the people that use the language. Mapundugun is spoken in both Argentina and Chile and here the government of a nation and the most powerful company of the world seem to be taken on because they do not represent the Mapuche people and are denied the right to decide for them.

In that light it makes perfect sense to go to court and insist that this is very much not wanted.

Thanks,
GerardM

Friday, November 17, 2006

To the winner go all the spoils

When changing you name makes you money, would you do it? Would you change your name for a pig or a goat ? Accepting such an offer could make me "Pig Meijssen", it would go well with my mascot. The name that the people changed their name to, "Hornsleth", is perfectly honourable. This "project" by the Danish artist Kristian Von Hornsleth is very much intended to demonstrate that international aid fails people. It fails people because it assumes that "the way of the donor" is best.

If you want to be helped you have to do this, that whatever. For me the great thing about Wikipedia is that it helps. It brings information to people. Its intention have always been to bring people information in their language. This means that culture, people and language are respected and that have people are enabled to help themselves.

For the western languages there is an abundance of great information. For many other languages, Wikipedia still has to take root. As the world becomes more wired, I expect Wikipedia will take root and become relevant for as a resource for both the culture, the people and the language.

When it is considered acceptable to bring information only in English, French, Arab, Chinese or whatever is considered a big language, the notion of Neutral Point of View that the Wikimedia Foundation offers is deminished. A NPOV exists because all points can be brought in the diversity that are the languages and cultures that are reflected in the 250 Wikipedias that currently exist.

It may be efficient to concentrate on the "important" languages.. but I do not want to consider the loss.

Thanks,
GerardM

Sunday, November 12, 2006

The worth of MediaWiki

According to the dark art of economics, everything can be valued. Everything can be given a price tag. People may object to this, and do on principle, but sorry, they have done just this for MediaWiki.

MediaWiki has a worth of $3.810.127,- when you assume that a developer costs $55.000,- a year and some other stuff. It was "valued" per the first of January 2006 and as the year is almost gone, it will be worth a lot more.

Another economic truism is that money makes money. Because of the success of MediaWiki more people will start developing MediaWiki. I am not in a position to deny this. WiktionaryZ extends MediaWiki. With new functionality much more becomes possible.

There is one thing missing in this argument. What is wrong is that MediaWiki is a tool. A tool that produces something that is far more valuable. Wikipedia is not the only project that MediaWiki enabled. I am sure some people who understand this dark art of valuation will be able to come up with a better number.

Given that money makes money, it is possible to leverage the MediaWiki generated content. Much proprietary content is not really relevant because it does not get exposure. By making content available it can get exposure. Material was often created to get exposure. By keeping it proprietary, thereby hidden from view, it does not do all that it could do. By making it available under a free license new opportunities arise.

MediaWiki enabled among others Wikipedia. WiktionaryZ has potential. My hunch is that like MediaWiki it will enable the creation of content in a different way. I hope and expect that it will help us to negotiate the release of much content under a Free/Open license and allow us to collaborate with many organisations and people.

Thanks,
GerardM

Friday, November 10, 2006

Running the interwiki bot for Wiktionary

I run the interwiki functionality of pywikipedia bot on all the Wiktionaries. It is a thing that I started and it is the kind of public service that needs doing. It links all the words that are spelled the same by adding "interwiki" links. These are the things that you see at the left hand side where it is indicated that there is also information in another language.

I have done this now for over a year and what I just noticed is the amount of words that I do not understand is growing rapidly. On the one hand it is to be expected as it is in line with the rapid growth of projects like the Vietnamese wiktionary. What now starts to happen more is that multiple wiktionaries have words together. That is what I see when I watch the bot.

In a way it would be fun to have WiktionaryZ in there. Currently we have 159.004 Expressions and we have 10.557 DefinedMeanings. Based on the expressions we would be the fourth project in size, it would be more reasonable to use the DefinedMeaning for the comparison and this would have us as the 26th in size.

Comparing Wiktionary with WiktionaryZ is like apples and oranges. Where Wiktionary has each word only once, WiktionaryZ counts them as existing in a language. Where there can be many red-linked articles on a Wiktionary page, the WiktionaryZ expressions are implicitly there.

It makes better sense to appreciate what the implications are of the numbers. In lexicology size counts. Only when people have a good chance of finding the information they are looking for will they find a resource useful. It is one reason why it makes sense to concentrate on certain topics or domains. WiktionaryZ is rich in ecological terminology due to the information that we got by including the GEMET thesaurus. By working on the OLPC children's dictionary we get a lot of the basic stuff that is the bread and butter of dictionaries.

Tuesday, November 07, 2006

Some thoughts on Alexa

Alexa is a website that provides an indication of the popularity of websites on the Internet. What I do does not matter, as Alexa only measures the use of Internet Explorer which is statistically becoming a less brilliant idea. Given the amount of people using Firefox on WiktionaryZ, I am sure that they do not know where many of the alpha crowd hangs out.

We have had our downtime this week and, I have been looking at how this affects our standing at Alexa. Sure enough, after two days of downtime, we hit the 900.000th place. Now that we are back up, we rebound nicely and today we are already back at number 478,316 for the weekly average.

WiktionaryZ has its own statistics, here you will find that our daily average hits did take a pounding. Given no more downtime and given that the trend of continued interest continues, the numbers will improve but the average will be depressed. At this moment all this is not crucial. When we get people to rely on WiktionaryZ, our service level needs to be much improved.

In a conversation with a developer I said once, when you respect our users, you have to treat them as if they cost us $150,- an hour. The point is that with the realisation how valuable contributors are, you are more likely to give them with the respect that they are due. With professional people using wikis, there actually is money paid for the time spend editing wikis this makes it more plain but it does not make a difference. Editors are to be respected and it is important to make the most out of what they do for us.

Thanks,
GerardM

Wednesday, November 01, 2006

The semantic web to the rescue ?

The reporting on the Internet Governance Forum in Athens is mighty interesting. Yet another nice article on the BBC website with some thought provoking ideas.

I find it really interesting that spoken languages are considered. However, practically at this stage the Internet is very much oriented towards written languages and, given the amount of stumbling blocks that exist to integrate languages other than the ones using a Latin script, I am afraid this is just a red herring.

I was also amused to see that the semantic web was brought to the fore as one solution to the problem of linguistic diversity. Yes, it is intended to be understood by computers. Computers are used by people and the semantic web is decidedly English. This raises the question how this computer that apparently understands English communicates to its user who does not.

There are however some great things to be said about the semantic web; first of all the terms used should be unambiguous. This in turn means that it should be possible to translate it to other languages than English. This is a challenge that we face in WiktionaryZ. We are able to have semantic relations and our semantic relations do translate to the language of the User Interface. So when the terms have been translated, in WiktionaryZ relations can be understood not only but also by computers.

This is a good moment for a disclaimer; WiktionaryZ is pre-alpha software. Many of the issues that have been tackled in the development of the semantic web we have not considered let alone touched. We hope / expect that we will be allowed to stand on "the shoulders of giants".

Thanks,
GerardM

Is this a "good" word

MediaWiki is the software that drives Wikipedia but also WiktionaryZ. It is probably one of the best pieces of software when it comes to internationalisation and localisation. This is demonstrated by the many localisations that have been done already for the software. Singing the praises of MediaWiki from me can be expected; why else develop WiktionaryZ on top of MediaWiki?

This does however not mean that all is well. I have written before about the problems with the Neapolitan language and there issue with the '' combination. Today I learned that a language called Hai||om uses the "pipe" character and consequently, I cannot make it work properly in a MediaWiki installation.

There are ways around such a problem; I can use one of the alternate names; San or Saan. I can expect that there will be no Wikipedia created in this language (only 16.000 speakers). But the point is, that even a system that does really well is only as good as the next language that proves that it has an issue with it's presumptions.

Thanks,
GerardM

Tuesday, October 31, 2006

The importance of good standards

At the Internet Governance Forum Mr Vint Cerf has said that changing the way the Internet works to accommodate a multi-lingual Internet raises concerns. The question raised in the BBC article is whether it is a technical issue or not.

Interoperability on the Internet is possible when the standards used are such that interoperability is possible. The current URL system is based on the Latin script. This was a sensible choice in the days when computing was developed in America. It made sense in a world when the only script supported by all computers was the Latin script. In these days computers all support UTF-8. All modern computers can support any script out of the box. This means that all computers are inherently able to display all characters. This does however not mean that all computers are able to display all scripts; even my computer does not support all scripts and I have spend considerable time adding all kinds of scripts to my operating system.

The next issue is for content on the Internet to be properly indicated as to what language they are. Here there is a big technical issue. The issue is that the standards only acknowledge the existence of a subset of languages. The result of this is that it is not possible to indicate using the existing standards what language any text is in.

Yes, the net will fragment in parts what will be "seen" by some and not "seen" by others. This is however not necessarily because of technical restrictions but much more because the people involved and the services involved do not support what is in the other script, the other language. When I for instance ask Google to find лошадь or paard, I get completely different results even though I am asking both times for information about the Equus caballus. In essence this split of the Internet already exists. The question seems to me to be much more about how to make a system that is interoperable.

The Internet is interoperable because of the standards that underlay it. With the emancipation of the Internet users outside of it's original area, these standards have to become usable for both the users of the Latin, Cyrillic, Arabic, Han and other scripts. It seems to me that at the core of this technical problem is the fact that the current standards are completely Latin oriented and also truly focused on what used to be good. At this moment the codes that are used are considered to be human readable. I would argue that this is increasingly not the case as many of these codes are only there for computers to use. When this becomes accepted fact, it will be less relevant what these codes look like because their relevance will be in them being unambiguous.

For those who have read this blog before, it will be no surprise that the current lack of support of ISO-639-3 for language names is one of my hobby horses. As I have covered this subject before I will not do this again. What I do want to point out that insisting on "backwards compatibility" is more likely to break the current mould of what is the Internet than preserve it.

Thanks,
GerardM

Sunday, October 29, 2006

New functionality of Firefox and blogging

I use Firefox as my browser. I have upgraded it to the latest version and now, my English will be spell checked for me in a real time fashion. The thing I probably will like best is, that a word like "localising" is now seen as correctly spelled. This is because Firefox allows me my British English. Blogger, although nice expects people to use American English spelling. This is useless when I were to blog in any other language.

Another thing that is nice is, that it allows me to accept words that are correct in the texts that I write; WiktionaryZ is such a word .. and so are MediaWiki, Wikimedia and Wikipedia.

By having spell checking done client side, the server functionality becomes cheaper for the service provider as well. The quality goes up.. all in all a good reason to upgrade to the latest Firefox.

Thanks,
GerardM

Monday, October 23, 2006

The NPOV of language names

Yesterday Sannab pointed me to this posting on the linguistlist. The gist was that there had not been full consultation with the academic community about the adoption of the Ethnologue database for the ISO-639-3 codes of languages. A secondary argument was that Ethnologue is primarily a religious organization and the question was raised if it could be ethically to have such an organization be the guardian of what is to be considered a language.

This e-mail is a reaction to what Dr. Hein van der Voort wrote in the SSILA-Bulletin number 242 of August 22 of 2006.

The problem I see with the stance taken is not so much in the realization that some of the Ethnologue information needs to be curated, it is also not in the fact that some would consider Ethnologue to be the wrong organization to play this part, the problem is that no viable solution is offered. The need for the ISO-639-3 list is not only to identify what languages there from a linguistic point of view, it very much addresses the urgent need to identify text on the Internet as being in a specific language.

At WiktionaryZ we are creating language portals. These language portals are linked into country portals, both countries and languages have ISO-codes. When I had questions about language names, Ethnologue was really interested in learning what I had to say about what are to me obscure languages. The point here is, Ethnologue wants to cooperate. Some people do not want to cooperate for ideological reasons and at the same time do not provide a viable alternative. This is from my point of view really horrible. The need for codes that are more or less usable is expanding with Internet time and not with the glacial time that is the time of academics.

When WiktionaryZ proves itself and becomes a relevant resource for an increasing number of languages, all kinds of services will be build on the basis of it using standardized identification for content. The ISO-639-2 code is inadequate. It is not realistic to expect ISO to review it's decision at this stage and not include Ethnologue. It is not realistic to expect such a review without providing an alternative that is clearly superior to what is ISO-639-3. It is clearly better to improve together on what is arguably in need of improvement than not to provide the tools to work with in the first place.

PLEASE COLLABORATE ..

Thanks,
GerardM

Wednesday, October 18, 2006

Learning a language .. because you must

When you want to go live in another country, to live there permanently, it is best to know the country, the language. When you want to emigrate to the Netherlands, it is often required that you first pass a test that shows that you have some basic ability speaking Dutch. This test is to be passed while still abroad. This test can be taken at a Dutch embassy; both for the embassy and for the people who have to take this test, it is a logistical challenge.

The test is to find if the "A1-" level of comprehension exists. People need to be able to listen to someone who speaks Dutch SLLOOWWWLY and uses a limited range of words. This list is finite. It is likely that many of these words are the same words that WiktionaryZ needs for it's OLPC project.

Many of the techniques that you would use in a school are the same as the ones needed to prepare for this exam. You need soundfiles, you may want illustrations; both pictures and clips, you want definitions in the many languages that people understand.

Given a list of words expected to be known for the "A1-" exam, it would be easy to get the communities of the would be emigrant to add translations for both the word and the definition. As there is a group of people that do not read or write for both, a soundfile needs to be produced as well.

The next thing is making this content available on the Internet and serve it as a public service. Maybe there would even be public money to make this a public service. Then again, even if there is no money available for it. It is a nice challenge and, when you can do this.. There are other countries that people emigrate to, even people from the Netherlands :)

Thanks,
GerardM

Monday, October 16, 2006

How to integrate wordlists in WiktionaryZ

There are many GREAT resources on the Internet. One I (re)discovered the other day is http://dicts.info. It provides information for some 79 languages and the information they provide is Freely licensed; the data can be downloaded for personal use as it can change rapidly.

So how do we integrate such information in WiktionaryZ ? WiktionaryZ insists on the concept of the DefinedMeaning. As this is central to how WiktionaryZ works, it is crucial that we have the concept defined. The dicts.info is split into two parts; a from part and a to part. The translations include synonyms and alternate spellings.

An application that is to include these translations could work like this: When an Expression is found in the to language and the translation is not there already, a user is shown the WiktionaryZ content with the suggestion to add the translation. This way the new information is integrated into WiktionaryZ.

The one part of this "how to" needed is possibly some discussion on the finer details, but certainly someone who will take up this challenge and develop this for us.

Thanks,
GerardM

Saturday, October 14, 2006

Eating your own dogfood

WiktionaryZ is about lexicology, terminology and ontology. Consequently you would expect that these concepts are in there .. Now they are.

For me not understanding a word and looking up what it means is a moral obligation to include it in WiktionaryZ .. the latest word I did not know was divot. It was used on the BBC news website and yes, you would appreciate what it would be like. The word in Dutch still escapes me :)

Thanks,
GerardM

Wednesday, October 11, 2006

Medical terminology

Yesterday I read an article on the BBC news website claiming that the term schizophrenia is invalid. The article points out that schizophrenia is not a single but a multitude of syndromes. The problem with the word is that given that it is understood to be a single syndrome, many patients are treated in a "one cure fits all" fashion. This is tragic as schizophrenia is seen as something that cannot be treated which is not necessarily true.

As we do not have much medical data yet, I have added the word to WiktionaryZ. WiktionaryZ will soon include an important resource of medical data, the UMLS. This will certainly include words like schizophrenia. I have added the resource that I created the definition on. With such a large body of medical terminology, it can be expected that many people will find their way to WiktionaryZ to learn what this medical terminology is. These people in turn will be interested in creating definitions that are consistent with what the scientific field considers something to be.

The word schizophrenia is well entrenched, it has a specific meaning to the laymen and my understanding from the BBC article is, that many professionals also need to learn about the ambiguity of the term. The question is, when is there enough support to depreciate the old meaning from a well known word like schizophrenia.. How will peer review work out in a wiki and how will the "public" react to these definitions..

Thanks,
GerardM

Tuesday, October 10, 2006

Localization

Yesterday I met some people at the Universiteit of Amsterdam the topic was there use of a content management tool called Sakai. This tool is to be used for a collaboration and learning environment for education. It is a tool that is used very much in universities.

At the UvA they want to use it as a shared environment for Dutch and Iranian students. This is in and of itself a splendid idea. The software is open source so much can be done with the software. I recommended that they should localize the software into Persian in order to provide a friendly environment. For me one of the bigger challenges was that Sakai provides a wicks environment but with a high level of authorization of what people can and cannot see. The challenge is how to make such an environment get to its tipping points where its community takes off and becomes autonomous in its action. Carving it up makes it much more problematic I expect.

The people who wrote the software however were braindead when it came to the use of their software in other languages. I learned from the University of Bamberg that they will not use it because localization is done by changing texts that can be found in the source code. A university in Spain I was told decided not to upgrade the software because it was not feasible to localize the software again..

It is sad when a tool with such promise is dead in the water because it was not considered that in order to be useful you have to allow for proper localization.

Thanks,
GerardM

Monday, October 09, 2006

Bi-lingual content

We are working hard to help the OLPC with content. WiktionaryZ seems to be really user oriented. WiktionaryZ wants to work together with organizations, with professionals. Quite often they have rich resources that are extremely useful, but need integration. Often this information comes in the format of a spreadsheet. When these are two column affairs with words in one language and words in another, it is not known what meaning these words have and, yes they can be imported but only people who know both languages can integrate it. At this moment it is manual work.

Manual work is sometimes necessary but it is time consuming and as it is, we do not have the tools to make it less time consuming. When such lists are imported with a "flag", it would be possible to locate them easily and compare them to existing Expressions for that language. This would help integration. When the Expression does not exist for both linked Expressions, it does not follow that the DefinedMeaning does not exist.

I am sure that we will have to deal often with these situations. We are at the stage where it makes sense to think about this. We are approaching the point where we have to deal with this.

Thanks,
GerardM

Friday, October 06, 2006

Using Microsoft as an initial standard

I do confess that I use Microsoft operating system and software. I will not actively buy new Microsoft software and I will not buy any new MS software if I can help it. Having said all that, it is clear that Microsoft's monopoly does not allow me to buy a laptop without paying the Microsoft tax.

I was using Word on my computer and I wanted to change the language as I learned not to write American but British English. Changing the language for a text I found a rich resource of languages that Microsoft supports. Of particular interest for me was how it splits English in many versions. This is something that I can easily emulate in WiktionaryZ. The question is; I can but should I. Would it be better to wait until someone wants to include something that is for instance Jamaican English or is it better to be proactive ?

Thanks,
GerardM

Thursday, October 05, 2006

Admin rights

At WiktionaryZ, an admin is someone that we trust to edit. When this someone edits, it helps us that he is a sysop. It allows him to block and delete errors when this is needed.

For those of you who know the Wikimedia Foundation projects, how would one of their communities react when a "bureaucrat" starts promoting users.. and create some 20 new sysops in an hour .. I am sure that it would create an uproar. Not on WiktionaryZ I am happy to say.
Thanks,
GerardM

Saturday, September 30, 2006

Bots

I run the pywikipedia bot for the Wiktionary projects. I have done this for quite some time, and what I do is a "public service".. The software is quirky; when it works, it works well. That is until recently when it decided to blank pages for no reason.

This was a great moment to update the software. This did not work; authentication problems. Sourceforge decided to have me change my password. Thank you sourceforge. This was the moment were it still did not work.

I asked Andre Engels to have a look. The result; the bot works again after some major chirurgy. Also the way it works for me is different; it assumes that I have a user on any wiktionary.. This was already more or less the case. It will now test more systems than before to see if an expression exists there..

All in all, I hope / expect that this solves my problems running the pywikipedia bot.

Thanks,
GerardM

The patron saint for the translators

Today, the 30th of September it is the day of St Jerome. He is best known as the translator of the bible from Greek and Hebrew into Latin. He is recognized by the Vatican as a Doctor of the Church. It is also the "International Translation Day".

WiktionaryZ will be a tool for everyone, it welcomes people from all countries and languages. Typically there is not that much political or religious to be found. This does not mean that we should not take a moments notice; we had Ramadan as the word of the day, and today I blog about St Jerome.

Thanks,
GerardM

Thursday, September 28, 2006

European day of languages

On the 26th of September, the European day of languages was held. We did not know.. This means that we are either not into languages or there was not that much marketing for this event.

The European Centre for Modern Languages has a website, it used to have posters agendas and all kinds of information about this. The stuff just does not load on my computer. It may be that the moment has gone, on the other hand many of the other pages of their website do not load for me.

I really wonder, if you have a website that does not work for people, does it serve it's purpose ?? Anyway there is always next year .. the 26th of September :)

Thanks,
GerardM

Monday, September 25, 2006

Cooperation ? Not with Debian...

I learned how "nice" it is to be "Free" at any price. Firefox, my favorite browser is Open Source. A great amount of effort was spent in making Firefox popular; advertisements in the New York Times covering a whole page.. The marketing of Firefox is truly one of the success stories of the Open Source world. Firefox is a trademark, it has a slick logo and the Mozilla Foundation protects its assets.

The Debian distribution is one of those "Free" distributions that does not appreciate that derivations of logos are not permitted. As it does not allow these files in it's distribution. Firefox insists that the logo and the name go together.. this results in a mess where Debian is likely to rename Firefox..

In my mind this is foolishness. Logos and trademarks are there to help products like Firefox to create a market. The evident denial IMHO of this is as stupid as the Wikimedia Foundation only having it's logos in the Commons repository and denying the logos of other organizations.

It is much better to appreciate that logos have a special place and, that it is great when organizations allow the use of their logo to be used with encyclopedic content.. For now is not to be because it is considered not to be "Free".. in effect denying this Freedom to others that is reserved for the own organization. It is not consistent .. it is a muddle.. It is bad practice.

Thanks,
GerardM

Sunday, September 24, 2006

When Mohammed doesn't come to the mountain ..

Some things are inevetitable; WiktionaryZ tried always to be a project about content. Getting the attention of people has been difficult because what is WZ about, what is it's relation with the Wiktionary projects, what is the relation with it's partners, paying developers and last but not least, what is the relation between WiktionaryZ and the Wikimedia Foundation.

At some stage, the "Special Projects Committee" of the WMF issued a resolution that they want to host WiktionaryZ. Combine this with the wish of Jimmy Wales for having WiktionaryZ in the WMF; it put us under some pressure. On the other hand, the experience with the InstantCommons project taught us that even friends need contracts to stay friends and it also taught us that when you want to get things done, it is often best to do it yourself.

With the election of Erik to the board of the Wikimedia Foundation, it seems that the mountain has come to Mohammed. Erik has been and is a key contributor to WiktionaryZ. This should facilitate a great relation between the WMF and the WiktionaryZ community and partners.

We have invested a lot of ourselves in WiktionaryZ. We intent to invest even more in the success of WiktionaryZ; we have schemes how it should integrate with the WMF projects. We see a bright future but as guardians of the project we will protect what WiktionaryZ stands for and nurture what we think WiktionaryZ will make possible.

Thanks,
GerardM

Thursday, September 14, 2006

Boinc

Boinc or the Berkeley Open Infrastructure for Network Computing is something I learned about the other day on a BBC worldservice program. I was fascinated by the wealth of options that it provides. It is quiet similar to the "Ligandfit" application that I have been running for the last 2 year and 153 days. It is different in that there is a much bigger array of things you can work on. Both have projects that I do recommend.

Given that many people like myself have loads of computer cycles going to waste, it makes sense to do something with it. For me there are two things that I do with my excess computing power; I have bots running doing maintenance work on Wiktionary and the Ligandfit..

Do you have cycles to donate ?
Thanks,
GerardM

Wednesday, September 13, 2006

Transliteration, do we need to do things twice?

At WiktionaryZ we support all the scripts UNICODE supports. This means that we already do both Serbian in the Cyrillic and the Latin script.. We also support Mandarin in both the simplified and the traditional script.. We want to support Cherokee which is available in the Cherokee and the Latin script ..

Many of these conversions can be done by a program or, like for Mandarin there are databases with both versions of the script. We received from Jeffrey V. Merkey permission to use a long list of Cherokee words, we also received permission to use a program called chr2syl that does transliteration automatically. There is a similar program for Serbian.

WiktionaryZ is at a pre-alpha stage, so we do not even have the basic functionality available, but would it not be nice when we add a Serbian word we automagically get the other version ?

Thanks,
Gerard

Tuesday, September 12, 2006

Papiamento

Papiamento is a language with two orthographies. There is the Aruban and the Antillian version of this language. Today I have the possibility to add this language to WiktionaryZ. The problem is that I do not have a clue what the code for the language should look like.

Thanks,
GerardM

Sunday, September 10, 2006

Antonyms in WiktionaryZ

An antonym is "a word or phrase that has exactly or nearly exactly the opposite meaning to another word or phrase". In WiktionaryZ, meaning is associated with what is called a DefinedMeaning. A DM is different in the way people think about concepts in that it refers to both synonyms and translations when you approach a concept from an Expression.

The technical problem is, when implementing antonyms in WiktionaryZ, are they still antonyms. Traditionally antonymy is considered within one language but because of the way WZ implements things, this is not true with in WZ.

Antonymy has it's own problems too. When a concept is considered the opposite of, this notion is often cultural. The problem then is, does antonym translate in the first place.

Thanks,
GerardM

Thursday, September 07, 2006

Internationalisation or internationalization or I18N

The interest that we have in localisation and internationalisation is well known by many, we need it for WiktionaryZ as we want a User Interface in the languages that we want to support. We want to support all languages.

One of us was approached by a really interesting company, a company asking us if we know people who are in the business of I18N. Asking us if we know people who are interested in a job doing a great job for a great company.

I find it really funny that we are seen as being (becoming) relevant for this subject...

Thanks,
GerardM

Wednesday, September 06, 2006

A small annoyance

The user interface is important, it is how people learn to work with a software environment. For WiktionaryZ the content is in the "WiktionaryZ" namespace. This is already counter intuitive as the default namespace is only supportive to what is done in WiktionaryZ.

The small annoyance is in the "New entries" it does show you only what is new in the "default namespace".. Not really relevant ..

Thanks,
GerardM