Monday, July 02, 2007

AGF

Yesterday there was an interesting meeting in the Netherlands billed as a "moderator workshop". There were many nl.wikipedians and some of them, myself included, were no nl.wikipedia moderator at all. To my delight this was deemed to be good; it was even considered that the name for this recurring event is wrong; in order to get more interest also outside the fairly limiting group of moderators, it is considered to have these meetings under a different name.

The most interesting presentation was the one where Siebrand informed about the state of play at Commons. I knew that Siebrand was one of those working hard to keep Commons clean and sane. It was another thing to hear him explain how these things are done... seeing the amount of pictures moved, the amount of duplicate pictures deleted. It is really impressive.

At the same meeting there was a guy who was articulating his opposition to Commons and its policies. He gave examples why he was disgusted with Commons. The issue as I understand it is in two things; the balance between the need for getting things done and the need for discussing individual issues and the understanding of the procedures used at Commons.

The need to get things done is easily explained. With 1.6 million images, and a growth that exceeds the English language Wikipedia, it is a big project. With all the pictures of all the different projects that should have their home at Commons, the amount of work that needs doing is mind boggling. On the other hand there are people who have a few pictures, who have strong feelings about their work and or who do not know and understand about the procedures at Commons. There is a balance between these two needs.

Over time, many changes have happened to the procedures at Commons, many things have been automated and have been improved to better reflect the needs that people have. If anything this is what I got from Siebrand's presentation. Siebrand indicated that there is still room for improvement and I got the distinct impression that a lot of work is done to do achieve even more.

AGF, Assume Good Faith, I do expect that the better nl.wikipedia moderator sign up to this notion. I was disappointed when it was not even acknowledged the improvements in Commons that were made. Well, that is probably the difference between a long serving moderator and a moderator that will help the community move forward in this brave new world where we cooperate and seek consensus on a continuously bigger scale.

Thanks,
GerardM

Tuesday, June 26, 2007

What language would be the one that provides the DefinedMeaning

At OmegaWiki we are of the opinion that the most valuable resource we have is the time of our contributors. It was what motivated to think beyond Wiktionary. We were wasting our time because we were not only interested in only one language. What we were working on was nice, but it was hard work to be relevant and we were with too few people to do the work well.

Now that we have OmegaWiki, we can enter data once and it is there for everyone. We now already have over 3.000 Expressions in Georgian, they are available to people in a Georgian, an English, a Dutch and many more interfaces. This makes OmegaWiki more relevant for Georgian then the Georgian Wiktionary; it has only 34 articles. We have some 11.000 Italian Expressions, this is roughly equivalent to the number of articles in the Italian Wiktionary that are actually about an Italian word. Our Expressions are available in all those user interfaces..

We are open to work together with other organisations, people. This is something that is also a standard practice in the Wiktionaries; much of the content is from different sources and is uploaded by bot. It is however sad that once the data has been imported, it is no longer possible to contribute back to the original source.

We are discussing how to collaborate with a great resource with more than 50.000 concepts in one of the major African languages with translations in English. This will get us enough concepts to have a relevant amount of content in two languages. The point is, they create their data in a completely different way; people suggest a word with a definition and an editor validates it and publishes it. So what could such a collaboration look like..

Well when our community adds content in this language, it can be seen as a suggestion to their editor. With new content accepted by the editor, it can be added to their database. Now here is the bit that some may find controversial; they have the copyright to their database and our choice of the CC-by license makes this collaboration possible. They retain their copyright, their license. This is not possible with the GFDL.

When we are going to import this data, what language would be the one that provides the defining Expression and the defining Definition.. I would say this African language :)

Thanks,
GerardM

Monday, June 25, 2007

Bolango anyone ??

Bolango is a language. It is spoken by some 20.000 people in Sulawesi, Indonesia. There is also a place called Bolango where some 5.000 people speak this language. That is to say, in 1981 there were some 5000 people who spoke that language in Bolango, how many people are still speaking Bolango I have no way of knowing.

Bolango is the 1000th entry on OmegaWiki of the ISO 639-3 collection. For most if not all the languages recognised by this standard we have a portal, there are over 7000 of these, so many more words will need to be added before this collection is complete. This notion of a work in progress makes it very much a wiki. As people work on particular content, OmegaWiki becomes more relevant for such a domain. Typically there are not that many words needed to cover the concepts associated with a specific subject and typically once these associated concepts have been defined texts about any of these concepts will not yield many more new terms.

Working on a list like the ISO 639-3 is different. It is a list and all these languages are quite distinct. It is particularly the language of linguists that connects any language with another. Having a collection like the one I am working on becomes relevant when people involved in languages find it relevant. Until that time, it seems like stamp collecting :)

Thanks,
GerardM

NB I just created a Bolango stub on Wikipedia

Sunday, June 24, 2007

Did I miss something

With increasing amusement I have been reading Kelly Martin's assessments of the people aspiring to become WMF board members. Many of the candidates have had the pleasure of Kelly's caustic pen. At some stage however, it got too much for me. Her latest had a few gems in it. Frieda, who happens to be the president of the Italian chapter, was attacked because a photo was showing some cleavage.. I had to go back and see.. did I miss something ??

Giving Frieda's position it is only natural that she wants to make the chapters more prominent. Kelly's assessment that the Italian chapter is run as a social club and chapters are for social networking.. Well, it is clear that Kelly needs some education on this. Given that she refers to Clay Shirkey's hell, I wonder what her role would be in such a social network.

In her "Nonbovine Ruminations" Erik Moeller was seen in at least as friendly a light. Erik took the trouble to answer; his conclusion after a lot of debunking: "Nothing I could say, no explanation I could offer, would allow you to rationally analyze what is really happening. You know all about me already, after all."

The high point comes in what Kelly has to say about Danny.. Here I feature as "Erik's trained attack dog". Well, I must show you sometime the certificate to prove the training I had to qualify.

Given all these pleasantries, there is a blog entry that explains why Kelly does not run herself: "it would only be good to get a great deal of attention (that is, drama)". Well, apparantly she has a craving for drama, it is the best way I can explain her ruminations. It also expect that this strenghtens her position in her home community.. in her own words: "I am widely disliked at the English Wikipedia".

The voting for three seats of the WMF board is imminent. I hope that the people who bother to vote take the time to read what board members do, what candidates have as their platform and particularly come to an understanding if the person they consider to vote for will have a positive and collaborative influence on our organisation. There is a lot of work to be done and when the board is burdened by people with a negative attitude, outlook it will make it that much harder to get something done.

Thanks,
GerardM

Friday, June 22, 2007

Godwin's law

The Polish want to get their way in the European Union. They are talking about the number of Poles that would exist when there would not have been the second world war. There would have been SO many more Poles...

Well, think of it this way, if there had not been a second world war, how many Germans would have been there... or Russians... or Jews.

Really, our Mr Godwin is right; do not use the Nazis, the second world war as an argument.. it only demonstrates that you failed.

Thanks,
GerardM

Wednesday, June 20, 2007

Logos dictionary

The Logos dictionary is one of my favourite resources on the Internet. I know the people well, I respect their work and their dedication, I even worked there for a month.

Logos has the ambition to make their resource the premier lexical resource on the Internet. They work hard on it and, I regret that we are not working together. We are not working together because the basis for cooperation needs to be very much a more wiki way of doing things. The software Logos uses does not provide a Wiki environment and the structure of the database does not allow for functionality that I would expect. This is one reason why we have persisted in developing OmegaWiki.

One issue that prevented cooperation was a license. the Logos dictionary did not have one, all it had was a copyright statement. This was fixed at the time, a GFDL license was added to the copyright. This means that all of the data was and has always been copyright Logos.

Sadly I became aware that Logos has removed the license information. Being the copyright holder to all the data they are allowed to do this. For collaborators and users of the Logos dictionary it is now less clear what the status is of this data. When you translate the Logos "Quote of the Day", are you entitled to attribution for instance ... ??

Facts cannot be copyrighted, collections of facts can be copyrighted. Really, without a license it is less clear what you can do with the data that is and has been for such a long time been freely available on the Internet. Be advised that this resource is older than any of the licenses that are currently so popular..

Really I love to collaborate with anyone but I would especially love to collaborate with Rodrigo, Cinzia, Gianni and Magdalena..

Thanks,
GerardM

Tuesday, June 19, 2007

Why I should join Citizendium

Today was yet another day to swell my head. I was invited to a discussion with one of the greats of the Internet. We talked for more than an hour and I was told that I was one of the thinkers of the Internet...

Thirty minutes later I received an e-mail from this amazing professional and academic organisation that started with "Dear Dr. Meijssen". In this e-mail I was urged to do something. Something I now have done...

I am part of an organisation that states that I am so really cool. Well, I work feverishly to achieve what this organisation wants to do. The other people in this organisation ARE doctors, professors, leading in professional organisations...

So Dr Sanger what do you say, should I join Citizendium ??

Thanks,
GerardM

Saturday, June 16, 2007

WIFI

Yesterday I had a meeting for the second time in the village of Lage Vuursche. The first time we were in a nice restaurant. It did not have WIFI for its customers. "There is no demand for this." I found this reaction astounding as I had just asked about it.

This time we were in the next door restaurant the "Kastanjehof". It did have WIFI; we had called prior to going there. We sat outside, it had stopped raining, we had our talk and our WIFI.

Thanks,
GerardM

Friday, June 15, 2007

Meta data for education

Wikipedia has proven itself to be really valuable for educational purposes. This is best illustrated by the support the Wikimedia Foundation gets from Kennisnet. When other countries would follow the example of the Netherlands, the WMF would not have a cash crunch.

Given that Wikipedia is particularly used by students, it makes sense to provide a better service to the educational process. The IEEE-LOM is a standard way of describing learning object meta data. Meta data can be associated with many of the resources that exists in the many projects of the WMF. There are two issues to consider; the data is relational in nature and the data has to be made available.

At the Holland Open Software Conference I met Erik Duval who is a professor from Leuven and who has a wealth of expertise on the IEEE LOM. He astounded me when he said that people should not enter the IEEE LOM data, automated processes should. Given that Erik is in the organisation behind the IEEE LOM standard, it means that for some of the dialects of this standard such processes must exist.

The data is relational. With the functionality that we created for OmegaWiki, we can have relational data natively in a MediaWiki environment. It is just a matter of associating this data with the articles that need tagging. This is not rocket science but the application of parts that already exist.

For the Wikimedia Foundation supporting educational meta data is a great method of making its content more relevant. The use that it will generate in education will provide a powerful argument why providing sustainability and investment in its organisation is in the interest of the national educational systems.

I invite the Wikimedia Foundation to consider this, not only as a method of raising funds, but first and foremost as a way to make good on its aim; to provide information to the world.

Thanks,
GerardM

Thursday, June 14, 2007

More Anthere at the Holland Open Software Conference

I have been on the lookout for the video registration of Anthere's presentation at the Holland Open Sofware Conference, it has not surfaced yet. What I did find was an exclusive interview that Anthere gave to the ANP. As ANP is a press agency, the same content could be found in several places, for instance here and here.

Several topics that were not covered in her speech can be found in the interview:
  • Anthere expressed the wish to adapt our software and improve the accessibility, the visually impaired are specifically mentioned
  • There is more about promoting our content in developing countries, not only do we want to make our content available as relevant is making their information available in our projects
  • With the growth of the Wikimedia staff, it is increasingly important for the WMF to be well organised and not only have finance be the top priority
  • Quality has a high priority particularly in the bigger projects
    • Self promotion is mentioned
    • The lack of sources is mentioned as a reason for deletion on some projects
    • Anthere indicates that she would not follow the Citizendium model where specialists check up on articles, they are fallible too.
Thanks,
GerardM

What do you do for the smaller projects

As I mentioned earlier, Anthere gave a great presentation at the Holland Open Software Conference. One question was asked and answered and it is still going around in my head. "What does the Wikimedia Foundation do for the smaller projects and languages ?" Anthere's answer was truthful; her answer was that we do not do much. When it does not happen now, we will wait for it to happen later.

It is a truthful answer and given the resources we have as an organisation, there is not much more that we can do. There are all the issues, all the dramas all the opportunities of the big projects to deal with. And these are to be dealt with either by the community of a project itself or by one of the staff of ten people that is to keep some of the biggest websites of the world going.

I discussed this with people like Sabine, and the conclusion is very much, there is no Wikipedia. There is an English language, a German language, a Neapolitan language etc Wikipedia. They are all individual projects. They may share many of the basic values, but in the end these communities, these projects are very much left to themselves.

So what do we do for the smaller projects. We very much want these projects to succeed. We now insist on some initial content and some initial localisation before we start a project in a new language but really once they are started they are on their own. There is no evaluation, no monitoring of the project and only when things are deemed to be REALLY problematic it may get attention.

So what should we do for the smaller projects. There are people, organisations who are willing to pay money for content in specific languages. This content can be truly in the spirit of the Wikipedia, it may be geared towards certain subjects. One of the best reasons for accepting and promoting this is that by creating a supply, a demand will follow.

Having created content on many Wikis, we have a grasp of what it takes to create content. The most relevant deliverable however is not the content, the hardest and most valuable deliverable is the creation of a community. This is hard because you cannot buy a community, this is hard because it is not clear what a community will consider to be important and this is hard because their opinion may not coincide with what is important to you or to an organisation that makes a content creation project possible.

We live in a world where deliverables need to be measurable. Content creation can be measured; you can pay a translator or a writer. You can spend money and deliver a product that includes interwiki, wiki links, images and conforms to style guides. But you can not guarantee that you build a community at the same time. Building a community takes time, it means that the people that make up this community need to be able to influence the process. All the right things can be done, but there is no guarantee that an autonomous community will evolve.

Anthere is right in many ways when she says we do not do much to help new projects. The Wikimedia Foundation cannot do much because it does not have the resources and it will happen; the smaller projects will take off. This process can be helped along by the creation of content. Growing the projects is a process, it takes people and effort. The growth of projects can be accelerated with investment however the mix has to be right to make a project truly part of the Wiki movement.

Thanks,
GerardM

Wednesday, June 13, 2007

The day after the Holland Open Software Conference

Anthere gave the keynote speech in Amsterdam at the Holland Open Software Conference. It was a great speech, it was so great because the speech was wholly dedicated to the Wikimedia organisation and its finances. This was possible because there was nobody at the conference that indicated not to know what Wikipedia was. For me there was little that was new, it was however great to hear it put coherently together. Many avenues of funding are closed to the Wikimedia Foundation because of the real or perceived problems that will result from within the community. Many avenues of funding are still very much closed because there is not sufficient staff to deal with acquiring funding. Anthere did a great job, it was a great opener because it put the sustainability of Open/Free projects as an issue very much on the map of the conference.

It was well published on the Dutch Wikimedia projects that the community was invited to attend the conference. It was cool to see Siebrand, Galwaygirl, Effeietsanders, Brabo. It is however a mystery to me where all the others were. This was a high grade conference with many speeches extremely relevant to what we do and there were so few of us.. A missed opportunity..

For me there were several new connections to make. I liked what Erik Duval said about forms to fill in for meta data like IEEE-LOM. He knows that much of this can be done in an automated way. I liked many of the things Massimo Mauro had to say about the European Union; I hope we will be able to cooperate in OmegaWiki. I loved the presentation of Eliane Metni, she connected the spirit of the open and free community with education in a compelling way; it is great that there is so much but it has to be adaptable in order to be useful in the setting of an educational process. A tool may be great but when it does not fit it is useless. There were many more but these are some of the highlights for me.

Next year the Holland Open? I will do my best to be there again :)

Thanks,
GerardM

Sunday, June 10, 2007

The point of OmegaWiki

Sabine wrote a great blog the other day about how combining content from Wikinews Positano news and OmegaWiki creates something that is really fascinating. Please read her blog first if you have not read it yet.

The great thing is that things have moved forward already. Martin Mai, who is with the University Bamberg, has written a first incarnation of software that assist the reading of an article on the Positano News. It provides you pertinent information one click away; it gives you translations and definitions it even allows you to go to for more details to OmegaWiki.

For me the most important point is that much more will be possible because OmegaWiki has its data in a database and not as a MediaWiki page. What Martin created is very much a mock up. It even works, but it is only a start. More is possible, it needs a fertile mind and some cool programming.

Thanks,
GerardM

Monday, June 04, 2007

Reputation

There is a long article in Informationweek about reputation. I read it with much interest. It covers most of the ground. The question that I do not find answered is, what does reputation buy me and, why would I care for an on line reputation.

My reputation is a consequence of the things that I have done. Some of the things I do have made me recognisable, I learned at a Wiki meet that my standard salutation, "Hoi" made me the first person that was recognised as an individual by someone who is working hard to understand the Wikimedia Foundation. This is my 250th entry in this blog, a growing group of people read this. I have written tons of articles in Wikipedia, Wiktionary and Commons and am now particularly active in OmegaWiki. But my reputation as I have it is a consequence of what I have done. It does not give me necessarily credibility; in Doctor Sanger's eyes it won't as he advocates certification in stead of reputation.

Much of the talk about reputations is defensive in nature; it is about vandalism, about pretensions, about why we should trust a resource and to what extend. This negative emphasis is self defeating because it does not value how a reputation helps in achieving goals. The cost of this absolute negative appreciation is that after a controversy a person like Essjay is no longer considered. He was once one of the most valuable Wikipedians and this was based on the good work work that could be observed.

The biggest problem that I can see with looking at reputation in a negative way is that you do not allow people to be wrong and the consequence is that this does frighten people off. People with a stellar reputation find it necessary to write under a pseudonym in Wikipedia because they are fearful of their reputation. When people are to be identified it does prevent people from contributing.

It is valuable to know what things are wrong. Scientific publications are a celebration of the positive discoveries. The dark side is that the discovery of things that are wrong is not as readily published, known. Many,many experiments are repeated over and over again because the fact and the proof why something is wrong is not published.

To me, a person gains his reputation by the work that he does. The more work done means the more opportunity for issues to arise it is however only the people that do that make mistakes. The people that allow themselves to be wrong should be celebrated. They are the giants on whose shoulders we can see further.

Thanks,
GerardM

Friday, June 01, 2007

A long tale ISO 3166-2:US

At OmegaWiki we are awaiting some new functionality.. Erik is away for a meeting and it will go life when he gets back. In the mean time I have been playing with a new collection; the ISO 3166-2:US is a list of US American administrative units including states and territories.

I was prompted by a guy who is really working hard to add vocabulary in Khmer. He had already done the countries and territories of this world and he followed that one up with this info. What I really liked was the distribution of the translations in other countries.. There was already quite a lot. There were the languages that were almost complete, while adding the collection I added the Dutch translations.

I hope people will find the collections a good start for adding translations in their language. Particularly the Swadesh list and the language lists are relevant. The language lists are particularly relevant because they are used in the user interface.

What you notice in all these collections is that there are some that are complete or almost complete and then there is a long tail of languages with some translations.. With time we will get more languages but we will also get more languages completed for the collections.

Thanks,
GerardM

Thursday, May 31, 2007

Uploading texts to Wikipedia

A friend of mine had worked on a text for Wikipedia. It was a great text and, I urged him to upload it. This did not happen. I tried to be polite and did not prod him aggressively. At some stage, it was almost a month later, I asked him about it.

His complaint was that it did not work. He had tried it many times only to find that the file was not one of those that could be uploaded. It finally dawned on me that my terminology was the problem; had I said that he had to copy and paste the article everything would have been fine.

Thanks,
GerardM

Sunday, May 27, 2007

A license for SignWriting

The SignWriting script is the product of a process that is already under way for over 30 years. It is a process that has been getting more and more momentum.

The script has been available on line for everyone to use from the start. When people have questions, need support, it has always been provided. People are learning to write their sign language all over the world. The last time I spoke to Valerie Sutton, the creator of the SignWriting script, it had to be short because people from Switzerland wanted to talk to her :)

As the SignWriting movement is ancient in Internet time, it started before licenses and copyright were considered. People needed to be able to use it, so it was made available. In a reply to a previous post on the subject of SignWriting, David Gerard asked about the license and the copyright and asked if there is an organisation behind SignWriting.

Yes, there is an organisation, the "Center for Sutton Movement Writing Inc." which is a USA nonprofit, tax-exempt educational membership organization. The copyright I understand is with Valerie Sutton and, she is considering what license to use. I have been discussing this with Valerie and the license that we are considering is the SIL Open Font License.

The big question she asked me to ask David and my readers is: "If all SignWriting symbols were under the SIL Open Font License (OFL) would you feel free to use the symbols?"

Thanks,
GerardM

Sunday, May 20, 2007

More on scripts, fonts and SignWriting

There were two bits of interesting news this week on scripts and fonts;
  • The Chinese script is much older, it is some 8000 years old. This makes the Chinese script some 3500 years older that previously thought.
  • Redhat made available the "Liberarion fonts", they are fonts that allow for the replacement of the proprietary Microsoft fonts. This is one of the impediments of adopting Linux. It is also one of the more visual aspects where Microsoft shows not to care about interoperability.
The Chinese script is old, it has had a long evolution and it is a living script. It is very much being used and one of the important aspects is that it brings together who speak different languages. The script is very much what unites the Chinese. It is not realistic to change Chinese for the Latin script because it, being based on sounds, will not serve in a same way.

The SignWriting script is young, it represents a revolution in the signing world and it has to be adopted by many signing communities. It is different from other written representations of signing languages because it is actually used for day to day use. As a script, it is different from Chinese because it represent movements like the Latin script represents sounds while a Chinese character represents in essence a concept. For this reason the SignWriting characters will mean different things in different sign languages.

The Liberation fonts make it possible to replace the Microsoft fonts without a need for reformatting the text. When you analyse this, it means that Unicode characters are now available in two interchangeable sets of fonts.

SignWriting does not have Unicode characters and it does not have fonts. The symbols that it uses are complete for most sign languages and some missing characters for the Ethiopian Sign Language are being added at the moment.

One really powerful argument why SignWriting is so important is because it helps deaf people to learn a written language that is foreign to them. English is in essence foreign to a person who grew up with American Sign Language; the written English language is not connected to the every day language of the student. When you learn a second language, you learn the shared concepts quickly. Now in order to know what these shared concepts are, you use a dictionary. Without a dictionary, without a written representation of the primary language, it is extremely hard to learn. It requires a really well trained memory.

It is exactly because of deaf people having to live in a world based on sounds that the bridge that the written word is so important. There is anecdotal evidence that kids who learned to write their sign language are better able to learn the written language of the spoken language that surrounds them.

It is for all these reasons that SignWriting emancipate the deaf. It will emancipate them because as a group they will become better able to communicate in their own world. A world that is both signing and speaking.

Thanks,
GerardM

Thursday, May 17, 2007

Dumb

Over the last month I have become interested in SignWriting. It is a fascinating subject for someone who is passionate about dictionaries. Given that OmegaWiki aims to include all words of all languages, it had bugged me for a long time how to include sign languages as well as written languages.

SignWriting is a recognised script, ISO-15924 Sgnw. It is only recently that it became possible to write a sign language grammatically correct using computer programs. It is written from top to bottom and it does not have its characters included in Unicode. There are some 30.000 characters at the moment and they are being converted from a bitmap into SVG.

I have been watching an instruction video on SignWriting, I have watched a video of some kids singing in sign language. When you see these people sign, I cannot even distinguish the individual signs, it goes way to quick for me. It is quite something, it makes me realise that I am dumb when it comes to signing and illiterate when it comes to writing sign languages. My redeeming quality would be that I am willing to be informed about it.

As the aim of OmegaWiki is to include all languages and as sign languages are as relevant as any other, I hope that the signing communities and particularly the SignWriting community will work with me to achieve this goal. Along this road there is the Unicode challenge and the challenge to get MediaWiki to support SignWriting.

This is likely to happen as SignWriting fulfils its promise and becomes the universally accepted script for sign languages.

Thanks,
GerardM

Tuesday, May 15, 2007

Wiktionary quality issues II

In a previous blog I wrote about the Russian Wiktionary being ostracised by the Polish Wiktionary. There was not enough content and consequently they did not want to have links on the Polish Wiktionary.

Yesterday, I blocked a bot run by an Arab Wiktionarian. I blocked it because it did not comply with the way interwiki links are created, it also did not have a bot flag. A link is only created when the words are exact matches. The problem at the Arab Wiktionary is that they have imported with a bot many words, English words, and they are all upper case.

I have blocked the bot because it is technically wrong. I do run my bot to correct the "damage". The biggest damage however is in the lack of communication between the Wiktionary projects. It is for this reason that I am of the opinion that though courageous efforts, many of them are failures.

Thanks,
GerardM

Saturday, May 12, 2007

Looking at Encyclopedia of Life differently

There has been a lot of comments on the Encyclopedia of Life in the Wiki community lately. All of them miss the most important point. The point is made by the Encyclopedia of Life itself; it will provide "aggregation and will use mash-up technology and Wiki style editing and accumulation of content".

Key in all this is that the Encyclopedia of life is not only Wiki style editing, it also provides for aggregation and mash up technology. It is the one thing that Wikipedia and Citizendium alike are incapable of. With the functionality MediaWiki provides, you are restricted to copying data into the article and having done that, it loses the connection with its origin in a practical way. This reduces these projects to sources of information because of the functionality.

When you look at the examples on the Encyclopedia of Life website, you find that they have information from many sources and present them together. In this way they provide a composite view on the subject that is considered. The data can remain fresh because once data is changed in the sources it consists of, this can be reflected because of the methodology.

One of the exiting things they provide in their mash up is the connection to old literature. This literature is extremely relevant to taxonomy as the oldest name that described a taxon is the one to be used to describe it. The consequence is that many names that were valid once are not valid anymore. It is interesting to learn how they will manage the linking of old valid names to the newer names. This has other applications as well, when an old paper mentions aspects of a species, it may mean that there is no mapping. It is then of relevance to consider the mapping to a subspecies or a different species that is referred to in the old literature.

Mapping the historical names of taxons in time is a hell of a job but an important job because it opens up the understanding of the old literature. For plants, the IPNI resource is of immense value. If anything, I hope that IPNI will be or become part of the mash-up.

Thanks,
GerardM

Tuesday, May 01, 2007

aderente

adarente is an Italian word. This word was defined in OmegaWiki. It has multiple meanings. One of these meanings, not the most common one, became a DefinedMeaning.

When you want to understand all this from what it says, it is extremely relevant to know that indeed "adarente" is the expression that is relevant to all this. When you do not, you will not appreciate that "conforme" is a synonym.

The whole notion of the DefinedMeaning is extremely important to OmegaWiki, it is fundamental. It is sad that, while we do record the word the Expression that goes with the DefinedMeaning, we still not make this visible.

Thanks,
GerardM

Sunday, April 29, 2007

Orientation of text

When you have a text, you write either in a right to left direction, like in Arab or Hebrew, or you write in a left to right direction, like in English or Russian. In OmegaWiki, we are going to support the orientation for right to left once Multilingual MediaWiki is integrated in our software.

I have given SignWriting a lot of thought lately. This is a script that goes top to bottom. What I just realised is, that as as consequence a left to right word mixed into such an environment would also go top to bottom. It makes however as much sense to have a right to left word also go in the same direction.

The problem is that when you have both RTL and LTR in a top down environment, it will seem odd.

Thanks,
GerardM

Friday, April 27, 2007

Google video

For a lark I checked Google Video for Wikipedia articles. It was no surprise to find stuff there. There was all kinds of stuff including this political rant about articles and NPOV on Wikipedia. The guy was insulting and threatening a Wikipedian in a way that I would ban him for a substantial amount of time. This is not freedom of speech is about.

I have no opinion at all about what this guy is on about. What I do know is that he does not help his case. This kind of gangster mentality is really obnoxious. His problem is that by behaving in this way he polarises to the extend that I would not even consider his arguments.

Thursday, April 26, 2007

Skype ..

I have been using Skype for so long now. It is absolutely essential to me. Today I had a first and, the person I spoke to also had a first in the way we were using Skype. To her Skype is an essential tool too.

I spoke with Valerie Sutton, her first was to use Skype and actually listen to it. Valerie is using Skype with video and uses sign language to communicate.

For me it was a first to actually see Valerie while communicating on Skype. It was all really accidental; I was about to disconnect when I found that the video was actually supported as well.

We discussed licenses.. I asked for some pictures to illustrate Wikipedia articles that are about sign languages. She told me that i could have any picture of her that I wanted and have it under a Free license.. We discussed the quality of pictures as well..

Now that the conversation is over, I thought of this, I would like to have a video fragment with every Wikipedia article about a sign language and have someone who is native in that language sign it :)

Thanks,
GerardM

How the Wiktionary interwiki bot is running

In a previous blog about Wiktionary quality issues, in a mail to the Wiktionary mailing list and a post on the English language bear parlour, I informed of the Polish Wiktionaries request / demand not to include the Russian and Vietnamese Wiktionary in the interwiki link.

The general reaction is that the Polish can wish what they want and that I should be willing to accommodate them. It is however not a zero sum game. Whatever I do, it has repercussions and it is in the end for me to make the choice that fits me, as the runner of the bot, best.

I have talked with Andre Engels about this. The result is that I will not update the Polish Wiktionary any more. All other Wiktionaries will still be updated with information with information about all Wiktionaries including the Polish.

I really dislike this situation. The good news is that the Vietnames Wiktionary is aware of this problem and is working on a solution, they just need time. The Russian Wiktionary is in my opinion a wasteland that is not really inviting to work on. Then again it is stubs on a grand scale.. people are working on it though. The Polish .. I am sorry that they feel this way. I do think that this only isolates them and that isolation is the major weakness of the Wiktionary projects.

Thanks,
GerardM

Wednesday, April 25, 2007

The lamest technology mascots ever

Wired has an article with some mascots. The idea was funny enough to get my attention. I got two things out of it. I did not know about the existence of Wikipede or that we might have had a mascot if it had not been for "a slow death by consensus". The other thing is the way our encyclopaedic project is described: " Gangware encyclopaedia Wikipedia".. Funny

Thanks,
     GerardM

Monday, April 23, 2007

Lies, damned lies and statistics

It has often been said that the only "reliable" statistics are the ones that you compile yourself. Strike that, I am not great at this science of statistics, I just use them to blind people with what is made obvious in this way. A week ago we celebrated that we had 15.000 Expressions for German. Today we celebrate that we have over 10.000 Expressions in Italian.

According to Kipcool, the statistics are wrong. His statistics take into account the fact that we occasionally have reason to delete Expressions. According to his figures we have 14.020 German and 9.776 Italian Expressions. This means that we may celebrate the round numbers again at a later date.. :)

For the Destinazione Italia project in OmegaWiki, we will need translations in many languages with an emphasis on European languages. As we already have a lot of content in Italian, it is important to mark what is already there also is known to be part of this collection. In this way do we know how many Expressions are there to process.

It seems to me that numbers, statistics tell a story, they tell as much about what is represented as about the person who uses the data. It is after the collaboration on statistics that you find what the numbers actually mean and what they truly represent. After such a process, they have become the statistics of a project and are owned by its community.

Thanks,
GerardM

Saturday, April 21, 2007

Wiktionary quality issues

On the Wiktionary project I run the interwiki bot. The process is simple; when an article exists in another language spelled exactly the same, I create an "interwiki" link. This allows you to see the information on another language Wiktionary. This process is an automated process, it works on all Wiktionaries and it is an unattended process.

I have received a request from the Polish Wiktionary to stop adding interwiki links for the Russian and for the Vietnamese Wiktionary. The reason given is one of quality. On the Russian Wiktionary many of the articles are created by a bot and they do not provide good information. An example is dispersion, there is nothing really in there. The Vietnamese Wiktionary is more problematic because a bot was used to generate declension and conjugation tables of Russian words and they got it wrong.

The Russian Wiktionary has some 81.000 empty shells and refuse to remove it. The Vietnamese are not willing to remove there incorrect data.

I have been asked to stop including the Russian Wiktionary and the Vietnamese Wiktionary when I run the interwiki process. To be honest, I run the bot as a service and I do not think it is the right thing to do. I think the Vietnamese are wrong not to correct the wrong data that they have. I am less sure about the Russian approach; in essence it is a stub. However, creating a Wiktionary in this way is like stamp collecting; you can look at it but there is not information about it.

Given how the process works, I am not sure that I can exclude either the Russian or the Vietnamese Wiktionary. The way it works is that I run explicitly on all Wiktionaries. When I exclude Russian or Vietnamese, I will probably end up removing all references to these projects. They are the third and fourth Wiktionary is size.

When I do not exclude the Russian and the Vietnamese Wiktionary, the bot may end up being blocked on the Polish Wiktionary. This will also kill off the interwiki process.

From my point of view, using bots to generate content in a Wiktionary only makes sense when there is at least a link to the word in the base language. When the initial creation of stubs is followed by the enrichment of these stubs it is acceptable. For having information that is completely wrong, there is no excuse.

The question is, will there be a discussion about acceptable practices in Wiktionary. The question are:
  • Can the Polish demand what they do?
  • Is having a project that consists mainly of stubs acceptable?
  • Is having incorrect data acceptable?

Thanks,
GerardM

Tuesday, April 17, 2007

The Uyghur‎ user interface

Uyghur‎ is a Turkic language spoken by the Uyghur people in Xinjiang. According to the English Wikipedia, the language is written in the Arabic, Roman and again the Arabic script. Cyrillic is actively used as well.

The officials script in China for this language is Arabic and has been since 1983. The ug.wikipedia has a Latin script user interface. The ug.wiktionary however is right to left. This means that there is at least a need for a user interface that is either in Arab and one in the Latin script. Given that Cyrillic is another actively used script, we need three message files for Uyghur.

When a language is expressed in so many ways, the current MediaWiki software does not allow for supporting Uyghur in a meaningful way. It will be good when Multilingual MediaWiki becomes a reality. The good news is that the prospects for this are really good. The last status report I heard was that some install routines have to be written and then people can start experimenting with this new functionality.

PS Check out what MLMW offers :)

Thanks,
GerardM


Friday, April 13, 2007

SignWriting

SignWriting is one way of expressing signed languages. There are several ways of doing this, signwriting is the one favoured by most people who actually write down the signed languages. There is a request for a Wikipedia for the American Sign Language or ASL expressed in SignWriting.

I am in favour of such a project. There are however a few relevant issues
  • SignWriting is currently not supported in UTF-8 and MediaWiki expects UTF-8.
  • There is software specific to SignWriting, it can be used for a Wikipedia
  • Due to the technical issues, there is no way it can use the Incubator.
Given the objective of the Wikimedia Foundation, supporting this project is very much something that we should support. ASL is a very different language. Given that the organisation behind SignWriting is very much in favour of the creation of a Wikipedia, there is ample scope for the Wikimedia Foundation and the SignWriting organisation to work together and overcome all the relevant issues.

For me it is important that with this Wikipedia project the culture of the deaf will get a boost. When SignWriting becomes even more mainstream, it has the potential to become the script for many of the other sign languages as well.

Practically, I would like to start with this project on WMF servers using the wiki like software that already exists. I would look for programmers and or funding to enable MediaWiki to support SignWriting. I would look for funding to create a full implementation of the SignWriting glyphs in UNICODE.

Thanks,
GerardM

Thursday, April 12, 2007

Braille

When you cannot see, resources like Wikipedia are not available to you. When you cannot see, you need something like braille to make such resources available to you. You still have the issue who is going to convert text to braille.

RoboBraille is an organization that converts to several different formats that can be used by braille readers that work with computers. The software works by receiving an e-mail with the text that needs conversion.

I can imagine that this engine could work for a website as well.. Consider what a difference it would make when Wikipedia would be available in this way.. The converted pages can be cached like any other page. When there is an issue with doing this realtime or near realtime, it would be possible to do this for articles that are featured articles or that are part of a "final version".

I would love to see a solution like this to become part of the service that we provide.. We aim to bring all information to all people.. blind people qualify :)

Thanks,
GerardM

Tuesday, April 10, 2007

Linguists go Wikipedia

In a previous post I mentioned that as part of the funding drive for the Linguist list, the subscribers of the list were asked to vote with their wallet. They did, and an intern will be paid to organise an editorial update of the Wikipedia pages on linguistics. The idea is that the many thousands of linguist that are subscribed to the linguist list will ensure that notable linguists like Eve Clark and Tanya Reinhart will get a mention.

Consider, this is a leading list of some 14,649 linguists that will be urged to help us improve both the quality and the quantity of the coverage of the field of linguistics.

I am really exited about this.

Thanks,
GerardM

Saturday, April 07, 2007

A board of trustees or an executive board

In a post on his blog Tawker suggests that it was "widely reported why" Mr Wool left his job. This is not true. Mr Wool explicitly refrained from explaining why he left his job and in this way prevented a discussion both about the organisation and his role in it.

The organisational consequences in the statements made in this blog are also very much wrong. The Wikimedia Foundation is in the process of setting up an organisation. This effort is hindered by a lack of funding and an unlucky choice in staff. As the staff becomes more professional, the board will be able to distance itself from the day to day affairs.

Thanks,
GerardM

Friday, April 06, 2007

Vietnamese

In OmegaWiki you can have many of the labels used in the data part in your language as well. For this to function, you have to select a language in the "user preferences" and there have to be translations of the words involved. Words like Estonian, Georgian or noun, verb, adjective will be shown in the selected language.

As the user interface is one of the most critical aspects for getting buy-in, it is something I often spend time on. Given that I do not speak languages like tiếng Việt, it is not always easy to find the right translations. For this language I was not able to find Gujarati. More bewildering was that I could not initially find the word Korean, it was there as korean with "tiếng Triều tiên" as its translation.

All in all I have added more than 10 translations in this latest session. I am sure that people who speak this language will have an easier time doing the same job.

Thanks,
GerardM

Thursday, April 05, 2007

{{WOTD|petition}}

At OmegaWiki we have a word of the day. This word of the day is hopefully created before a new day starts. It is something that I find myself doing almost everyday. Most often I just pick a word at random. Today I picked the word petition.

Today I signed a petition, the Alan Johnston Petition, at the BBC News website. I think it is really sad when journalists are the victims of political violence. It prevents news coming from corners of the world. In this way you can prevent news going our or coming in. By stopping the flow of information, the risk increases that the "other" will be seen as the enemy.

I am afraid the Palestinians are shooting themselves in the foot. You can also sign the petition.. or think some positive thoughts about this.

Thanks,
Gerard

Tuesday, April 03, 2007

Linguist List's Wikipedia Update Vote

Linguist list is the premier mailing list that aims to provide a forum where academic linguists can discuss linguistic issues and exchange linguistic information. It is more than just a mailing list; it also provides fellowships to graduate students who serve in return as editors on the list.

Yesterday there was a message that following an initiative on the Russian Wikipedia, they would pay for a graduate assistant to work half time for one semester to coordinate the improvement of the field of linguistics in the English language Wikipedia.

When their community thinks it important, they have to spend $2000,- extra and this will pay for this initiative.

There is nothing that stops people not yet associated with the Linguist list to contribute to this funding drive as well. :)

Thanks,
GerardM

Thursday, March 29, 2007

Scripts on the Internet

I am back from the ICANN conference in Lisbon and, I have the T-shirt to prove it :) At a meeting the BSI proposed to create a standard that will describe how a mandated list with a ccTLD for every country or territory is to be produced in other scripts. This would be an industry standard that would eventually be adopted by ISO. Having such a list would be truly beneficial to get to the stage where the registrars for these registries know what codes to use. I have been told that a list has already been compiled for the Arab and the Cyrillic script.

On the BBC-news website there was a great story that explains why it is so important to allow for content in the language that people speak ..

One other thing I learned is that you can already have a .org domain name that is in an UTF-8 script. To me this is quite important because it proves that supporting UTF-8 is something that can already be done. It would make such a difference if the Internet was as functional as it is for us people who read and write the Latin script.

Thanks,
GerardM

Monday, March 26, 2007

Cryptography

Today I am at the ICANN conference in Lisbon. For me it is quite special to be here. One of the nice surprises was a gentleman editing on the Wikipedia article on Implicit Certificates. These are some special certificates for use in smaller devices. He mentioned that after writing this article, within the hour relevant edits were made to the topic.

The great thing for cryprographers is that Wikipedia provides both support for LaTeX and as relevantly, it can be updated when new developments happen. This makes Wikipedia an excellent environment to maintain information on this domain.

Thanks,
GerardM

Thursday, March 22, 2007

Citizendium and licensing

In yet another diatribe Dr Sanger informs us about the differences between Wikipedia and Citizendium.

The thing that struck me most are two things; contributors have to give a non-exclusive license to Citizendium AND the license is now to be the Creative Commons CC-by-nc license. The consequence is that only the Citizendium organisation can license commercial use, obviously for a price. They assume that as it is to be written by experts it will have value. It also means that once licensed, the licensee can do whatever.

When you compare Citizendium with Wikipedia, you have an English only versus a multi-lingual project. You have a project that covers almost everything and a project with a few thousand articles. You have a Free project and a project that is increasingly restrictive. You have a project that informs the world with NPOV information and a project that is to be written by experts.

Really I think Dr Sanger is doing a great job promoting Wikipedia by increasing the differences.

Thanks,
GerardM

Wednesday, March 21, 2007

Supporting the visually impaired

The content that is created in many Wikis is really relevant. Wikipedia for instance provides you with unsurpassed encyclopaedic type of information. This popularity can be deduced by Alexa's current ranking for today .. number 9. For us users and editors of Wikipedia it is like a roller-coaster; it is exhilarating :)

For people who are visually impaired, several of the Wikipedias have projects where people record the articles .. Really nice, really relevant.

Today I learned about new software promoted by UNESCO called Sakrament Libreader, it allows for text to speech conversion for English, Russian and Belarussian. Just consider, the English Wikipedia takes, according to Alexa, 53% of our traffic. There is a certain percentage people that are visually compared to the unimpaired.

Consider what would happen if the Wikimedia Foundation would support and promote this kind of functionality. Not only would many people be helped by this, it would also give an impetus to make this functionality available for other languages..

Thanks,
GerardM

Saturday, March 17, 2007

IATE became available

IATE or the Inter-Agency Terminology Exchange became available as a resource on the Internet. An introduction to IATE explains nicely what IATE is about; it is about the terminology of an organisation; the European Union. It aims to demystify the jargon used by the organisations of the EU.

In the past IATE was accessible as well; it proved popular and it was quickly hidden behind a password. This time you can get access by using the URL http://iate.europa.eu and it hopefully means that this resource is now officially available.

When you compare IATE to what it replaces, it is a massive step backwards from a copyright point of view. EURADICAUTOM was available under a much less restrictive license. It would be nice if the EU would steal a page out of the US book; most of the information provided by that government is available without restrictions.

Thanks,
GerardM


Tuesday, March 13, 2007

Google Summer of Code

Google has announced its third Google Summer of Code. This is an annual event where students develop on Open Source projects. This is definetly one of those activities that does a lot of good. It is one way whereby Google makes its mantra of "do no evil" work well.

For Open Progress, we have entered for a first time; we have a nice mix of MediaWiki and OmegaWiki based projects. All these projects are dear to us. We have shown Brion our list, and we are likely to work together on these.

What struck me is that when you apply for the GSOC, it is compulsory to have a mailing list. This is the traditional way of doing things. I am subscribed to many mailing lists. I think mailing lists suck big-time. There is so much repetition, the signal to noise ration is typically quite bad. I do not understand why people do not use a wiki to document and discuss.

I think this is one of those instances where software development proves to be conservative. When you follow the subjects you are interested in on a wiki, you can use watch lists to make a selection, you can use RSS to follow the changes on a low bandwidth wiki.

Because you refactor what is there, there is no need to repeat so much. When you have discussions that are getting out of hand, backrooms can be opened for those quarrelling. Maybe I am an idealist that I see it in this way .. oh well ..

Thanks,
GerardM

Sunday, March 11, 2007

Fon or Meraki .. I want the functionality of both !!

In the field of Internet connectivity, WIFI is the thing that keeps a road warrior and Internet junkie sane. It also keeps him poor. The amount of money some companies dare to charge for connectivity is tantamount to high-way robbery. Particularly at places where you have to spend lots of time, like hotels and airports are really unfriendly places. Particularly in airports it is galling; you have to be there so many hours in advance, you are often delayed .. it would make more sense to provide it for free and in that way keep the punters happy.

With many Internet organisations increasingly dictatorial in what you can and cannot do, with intellectual property organisations only interested in filling the pockets of the big companies, WIFI is a next place where the people can create a commons away from these established malpractices.

Fon is a WIFI sharing organisation where you provide access to other fonistas by sharing the Internet connection. In one of the more innovative approaches they are targeting the neighbours of Starbucks with free routers. The business model is that you can pay a small amount for access to the network.

Meraki is a WIFI sharing organisation where you provide Internet connection to an area using a mesh network. By including repeaters in strategic places, the area covered can be quite extended and, multiple Internet access points can be part of the same network. When you operate a network, you can determine who can access the network and even provide access for money.

I would like to have a mix of both. At home Meraki would be ideal because my ISP uses a different technology from the ISP of my neighbour. This means that I will improve the connectivity both for myself and for my neighbours. It also allows for providing truly local information. Fon would be ideal when I am away, with an increasing number of fonistas it means that me providing connectivity is the assumption of good faith that will find its sweet rewards.

Combining the two would be awesome. I would not hesitate long to go that way.

Thanks,
GerardM

Wednesday, March 07, 2007

Localisation of MediaWiki

There is a policy for new languages in the Wikimedia Foundation. One of the key things is that we want to prevent new abominations of projects where the language is not what is advertised. We have seen these in the past; one of the worst in this is what is called the Belarus Wikipedia; the people in control of this project prevent the use of the Belarus language as it is used in Belarus. This is a really awful situation and the language commission has asked the board repeatedly to act on this.

What we want to achieve is that new projects will promote cooperation in stead of establish division. Even though linguistically there is a substantial difference between the different forms of English, it is generally accepted that there will be only one English Wikipedia. This means that when the differences between two languages are less than for the different forms of English, the language committee is not likely to approve a new project.

When there is merit for a project in a new language, the people promoting this language have to show their commitment in the Incubator. This is where the environment in that language for the requested project is set up. What is expected, is that there will be a number of well written articles about different types of subjects, there should be a main page and, the most visible parts of the user interface should be localised.

The sad thing is that the aspiring projects cannot conform yet to the requirements; when a new Incubator project is set up, the message file for the new language is not created. What is needed is for one of the developers to create the necessary files so that the localisation can start.

For many languages in the past, there is no message files either; it means that the localisation is done locally and that this effort does not lead to the localisation of the MediaWiki software.
I am really pleased that for one language, Marathi, many of the messages have been imported into SVN by Nikerabbit. In a few days Marathi will be supported in all WMF projects. :)

I would welcome it when the Wikimedia Foundation gives the support of the minor projects a priority. At this moment it has none. The creation of message files in the Incubator for all languages and, when a language becomes a project, the inclusion of the first localisation into the MediaWiki software is the bare minimum. When we boast that we have localisations in some 250 languages, it should be a verifiable truth.

Thanks,
GerardM

About dictionary writing

Connel MacKenzie is one of the English language Wiktionarians who has had a big influence on the development of the Wiktionary project. He published his notions about what a dictionary should be. As I have posted a response to what Erin McKean, the current editor in chief for the NOAD, said in a presentation at Google, it is nice to write about in response to this as well.

For Connel, the project is about the English language. This is a big difference in approach to what Wiktionary is said to be about. He wants to limit it to those words that are part of the 600.000 most common terms. This is problematic because how do you judge something to be common and, what is common in one branch of the English language might not necessarily be common in another. He is of the opinion that "freak" terms should only be there in a sanitized form. To me it is important that a term is clearly and fully explained. When you "sanitize", it is not clear if the full meaning survives for someone who does not know the term. By disallowing multiple word entries, you loose the connection to those entries that are single entries in another language..

In his commentary, Connel writes about the restrictions that faces Wiktionary that are the consequence of its flat file format. You can not segregate different types of content when the basic technology does not support it. At that he would be better off being part of OmegaWiki as its technology allows for all the things he is looking for.

He hopes to get a useful Wiktionary when he has a dozen programmers available for such a project. At the same time he despairs because of "the current anarchy" it may take ten to twenty years..

I do admire the constructive work Connel has put into Wiktionary. I doubt that Wiktionary will ever become useful other than as a resource where you can look things up on the Internet.

Thanks,
GerardM

Tuesday, March 06, 2007

Upper ontologies

"An upper ontology attempts to create an ontology which describes very general concepts that are the same across all domains". As almost always there is a Wikipedia article about this. Given that the subject is difficult, there is since August 2006 a request to clean up this page .. I have to agree that the subject is interesting and potentially controversial.

For OmegaWiki, an upper ontology is one of those things.. An upper ontology defines the broad strokes, and while it descends downwards more and more aspects may be inherited from higher levels. This means that an upper ontology has practical implications. It also means that we will eventually have in effect an upper ontology by default.

As an upper ontology creates the concepts that are true across domains, Wikis for Professionals will want to define how their ontologies fit into the upper ontology. As they are bound to have overlaps with other domains, there will regularly be found to be in conflicts between domains. As one WfP needs to link into other domains, the question of primacy will raise its ugly head. One WfP was there first, but this other domain is not its competency... A new WfP does have the competency but is disrupts the existing WfP...

An initial selection of an upper ontology will be crucial, its evolution will be exceedingly important when OmegaWiki will is to be bound by the integration into it. Personally I expect that a different model will arise; one where on the one hand great care will be given to the evolution of the upper ontology while on the other hand functionality will be created irrespective of the upper ontology.

Consider, when you know that something is a plant, you can infer all kinds of things about it. It is not really necessary to know how it fits in the greater scheme of things. It would be nice, but it is not required. When a lot of functionality is determined in such a way, there will be the question how this will fit together; it is therefore my prediction that these two forces will eventually find a balance. As OmegaWiki matures, this balance will become increasingly stable.

Thanks,
GerardM

Wednesday, February 28, 2007

NIH Wiki Fair

The National Institutes of Health organised a Wiki fair today. There have been great presentations; the line-up of speakers and subjects was impressive, many of the presentations are on line and the whole presentation can be seen as a video as well ... all 5:41 hours off it...

For Wiki aficionados particularly interesting are a presentation of Dr Larry Sanger, and Dr Barend Mons. Larry is part of the Wikipedia history and Barend is part of the OmegaWiki presence.

Larry's presentation is self serving; he gives a selly presentation about Citizendium and he urges the people from the NIH to join in his project. That is in and off itself OK as he just what he has to do. As the raison d'être of Citizendium is in doing a better job than Wikipedia there are loads of comparisons with Wikipedia.. It is not surprising that you find no nicks when that is your policy, it can not be considered something that is "good" in comparison. Given that Larry finds it relevant to be known as Dr Sanger, I would expect a more scientific approach to his presentation. Given the small size of the project, it is not so strange to find that there is a lot of harmony around. This is very much the experience of all Wiki projects. It has been well documented in many scientific papers and, I am sure is Dr Sanger aware of this.

Being part of the history of Wikipedia and having created Citizendium as a reaction to what is considered "wrong" in Wikipedia means that Citizendium will always be compared to Wikipedia. I think that having started from scratch is daring.. It is the right thing to do. There are certainly more scientists working on Wikipedia, and given that Wikipedia has more than 1000 times more articles, it will be interesting to see what kind of attention this project will get.

Thanks,
GerardM

Saturday, February 24, 2007

Open Access

Given that I am very much in the open source and open content world, it will not be a surprise that I am very much in favour of open access and that I do not think highly of the patent system.

I think that it is easy to argue that patents kill. My argument goes like this; the pharmaceutical industry is documented to be only interested in patentable medicine. When a medicine is not patentable, they do not have an incentive in researching its application for particular purposes. Recently the university of Alberta found by accident that dichloroacetate (DCA), appears to suppress the growth of cancer cells without affecting normal cells. The problem is that there is no funding for doing the research to prove the efficacy of this medicine. It is therefore easy to understand that only substances that are patentable are considered in much of the bio-medical research.

I think that is is easy to argue that the classic business model for scientific publishing kills. The cost of reading scientific articles is staggering. The article published in Nature about what is demoed on wikiprofessional.info, for me it is a really exciting article, for me this really relevant article costs $30,- to read. I can not lawfully send it to my mother to read. I can only tell her. Scientific journals assume that the only people who need to read their papers are scientists; the scientific libraries will have a subscription and that is how it has always been. Except that it is not true. Laymen have as much a need to read bio-medical science papers; they or their loved ones are affected by all kinds of afflictions. They have a need to understand what is happening to them. Often doctors do not know all the details that are available and , there are enough documented cases where research on the literature by laymen made a difference. One famous example is known as "Lorenzo's oil", a story about a disease called adrenoleukodystrophy or ALD.

Not only laymen are denied access to literature, many universities cannot afford the cost of literature. With the science results owned by business, it is important to understand how vital the whole "open" movement already is. A new, a different business model is needed and many organisations are developing these. The big challenge is for science to become science again; to be able to research without having the public pay again and again and some companies walking away with only an eye for profit.

The challenge will be to find a balance that does justice to all parties involved.

Thanks,
GerardM

Thursday, February 22, 2007

Mother language day

Yesterday, it was the International Mother Language Day. There is a program in Paris with all kinds of speeches and happenings. I read the program, I know some and I know about some of the people involved. I wish I was there. I wish they recorded some of the speeches, presentations.

One thing I find funny is that UNESCO indicates that there are some 6000 languages where SIL indicates the existence of over 7000 languages and where the ISO-639-6 will know at least 25.000 linguistic entities .. :) I wonder if they need to be emancipated so that UNESCO will feel a need to at least acknowledge their existence..

It would be cool if International Mother Language day had hit the Internet; a podcast would have been nice ..

Thanks,
GerardM

Sunday, February 18, 2007

A root canal treatment anyone ?

Yesterday, I had an appointment with a dentists specialised in the noble art of endodontics. I had a molar that had already had a root canal treatment before and it needed some more work. I was nervous. Actually I am terrified of dentists and consequently I am my own worst enemy as a person suffers most from the suffering that he fears and that often never materialises.

So I went to my appointment and I was wearing a Wikimania 2006 t-shirt. A person waiting in the reception reacted to this and, we got to talk about science, creating educational content in a collaborative way, licenses and copyright, OmegaWiki, the Nature article and the Wikiprofessional demo.

When it was my turn to sit in the "chair", the doctor had heard much of the conversation and asked several questions.. There was no time gap between being operated on and talking about the things that are so dear to me. One big difference between what an endodontist does and what a dentists does is in the tools of the trade; an endodontists uses microscopes and a lot of digital imagery. It really gave me the feeling that I was operated on. Anyway, after the operation I gave a demonstration to my endodontist.

The good news is; I did not have time to be nervous.. Two more people may have a good look at OmegaWiki. It feels good even though my molar is still sensitive .. :)

Thanks,
GerardM

Friday, February 16, 2007

A friend of my is a scientist...

A friend of my is a scientist. He is a terminologist. He has published a lot of papers and, he is considered one of the best in the field by some other people I know.

I told him about wikiprofessional, what kind of things we are doing in the bio-medical field. What we could do for other fields as well when we have the terminology available to us. This was two days ago. Today I learned that much of his work was once available on diskettes. These diskettes were no longer there .....

I discussed this with another friend. She has some great OCR-software. We know each other, so when the terminologist scans his paper paper, the translator can translate it from a analog into a digital format. The terminologist can define his terminology, his papers will be known but as relevant, we will have started supporting the terminology of terminology in OmegaWiki.

It is great to have friends ...

Thanks,
GerardM

Friday, February 09, 2007

Cleaning a bit on the sk.wiktionary

Many of the Wiktionary projects have a moribund existence. At some stage people worked on it. They left the project, things changed and where never properly taken care off. Today I was in an extended chat and I was not the lead talker, so I had some time to do some stuff that does not take much attention.

I cleaned up a bit on the sk.wiktionary. At some stage all Wiktionaries supported proper casing. This meant for the acronyms like ADHD that they became aDHD. Somebody needed to go and fix these so that they became properly ADHD. Today I fixed more than sixty acronyms that started with a, there are many more of these that need fixing and, not only but also on the Slovak Wiktionary.

I am really pleased that at OmegaWiki we only need to fix things once..

Thanks,
GerardM

Thursday, February 08, 2007

The Shtooka recorder

I was told to have a look at the Shtooka recorder. It is a tool that makes it really easy to record the pronunciation of words. You provide it with a list of words, you provide the meta data and, it creates an .ogg file for you in the directory of your choice. It makes it really easy..

For me the challenge was to understand how to use it. There is no user manual and mostly it is self evident. So when I understood that I had to press on the "record" button, it started to work.

This software does do wonders when recording words that are provided in a Latin script. I tried it to record a Persian word, کم حرف and it failed. I checked it with something Russian and it failed as well.

For me it is a big improvement, I did some 15.000 words with Audacity.. It would have saved me a lot of time when I would have had Shtooka.. The next thing to automate; the uploading to commons...

Thanks,
GerardM

Friday, February 02, 2007

Ten things you want to know about dictionaries

I met Erin McKean at the Wikimania 2006, I loved her presentation then and I was really happy when I found her presentation in the Google Author series titled "Ten things you want to know about dictionaries". I loved it, I have seen it twice now. I may even have added value to it by adding the "dictionary" label to the presentation. Here I am going to react it with my OmegaWiki hat on. So yes, please watch the presentation (almost an hour and well worth it) as I hope it will improve the understanding of my reaction.

There is no one dictionary; it is a tool
OmegaWiki as a resource is very much a child of the Internet; consequently it has the potential to configure its use. When people do not care for particular information; they should be able to make it invisible or turn it off. In a way it is like the cordless drill, by replacing the drill with a different thingie it becomes an other tool. The same is true for pronunciation; we love to include IPA, but we can also record pronunciations this way people do not need the understanding required when reading IPA.

Please read the "front matter"
People indeed assume that they understand tools like dictionaries and wikis for that matter not to RTFM. For a consumer good like a read only lexical resource, it is pretty safe when the introductions have not been read. As OmegaWiki allows people to add/edit to the information that is in there this proves to be much more problematic.

Inclusion in the dictionary is because it is useful
As we do not have all the functionality that we need to be a credible lexical resource, this is very much a state we hope to get at. However, our aim to include all the lexicological, terminological and ontological sounds pretty like megalomania. Our standard excuse is that we use is that is already less problematic because we only want to do this once and this is where we came from. The data will only be useful when there are people who care about particular categories of data. I am totally with Erin that only data that is useful should be included. Getting rid of unnecessary cruft is hard work.

Horrible words make it in their too
I am a fan of swear words in dictionaries, particularly when there is some etymology to it. Most often people use swear words as an expletive without much understanding for their actual original meaning. As English is for me a second language it is relevant for me to understand why I would rather be a bigot than a racist or someone who discriminates.

The other part of horrible words are those words that are actually used and offend the aesthetic sensitivities. In several medical resources you find stuff like MALARIA and Malaria. UGLY. However as it is useful to these folks, it makes sense to include them anyway. As they are exactly the same as the preferred English expression of malaria, it does not hurt.

Words like "irregardless" well being a non native to the language, I just want to be able to find them.

You have to look at all definitions to find the REAL meaning
The way the New Oxford American Dictionary does this is exquisite. They use a core sense / sub sense approach. To me this seems an approach that is very much language specific. For OmegaWiki to have such an approach, it will need quite a lot of thinking on how to build this.

Approaching the understanding of an expression with core senses / sub senses could be one way of stretching the number of concepts that people can juggle with. For Operational Definitions there is currently this practical limit of some 7 different meanings.

Dictionaries have a sell by date
For OmegaWiki this is not an issue as it is web based resource. However, the same issue still applies; a Dutch book printed in 2003 will use the orthography of 1995 and not the 2005 orthography. People will still want to be able to understand what this word means; annotating them as not being the official spelling since 2005 is relevant.

When words are tagged as used in earnest up to a certain date, I could even include 15th century German words and not have people be confused.. filtering would also help here ..

Facts are good
Referring to actual usage seems obvious, what we intend to do is link OmegaWiki's content to Wikipedia, this will be the most obvious resource to start of with; we aim to have a Wikipedia in all languages. The language is modern usage, so given our Wiki credentials it is the obvious corpus. I totally agree when it is said that we are limited in this way; it is however a great start. When we gain a community with people, organisations that introduce us to other resources, it will be great.

At this moment OmegaWiki is still very much like a "stamp collection"; it is a nice collection, and at some stage it will even become useful.

What we do is like an iceberg
The work done on the New Oxford American Dictionary may be in preparation for the moment when other information like thesaurus information will be included. For OmegaWiki, including information from thesauri is what we did from the start by including the GEMET data, OmegaWiki at this moment is very much "you get what you see", there is little of an iceberg yet. In a way this prevents the usefulness of our data because there is often too much to take in.

Etymology
Technically etymology is one of the hardest nuts to crack. When a word has its root in Latin, it often came to the English language through a French or Spanish connection. I wonder how the NOAD does this, indeed I do not have a copy so I have not read the "front matter" either :) .

Neologisms
At this moment there is not much of a problem about neologisms yet. Our community is still small and sane.. sort off (you must be weird to involve yourself in a project like this). My current thinking is that this problem can be solved using annotation. When a word is tagged as "Neologism; this word is not used except by the author" it will be pretty devastating to the prestige of the word and or the author on OmegaWiki.

The bonus: Using Google and other resources
Erin explains that the resources for building a resource like the NOAD are hardly as much as she would want. Given that this is true for a successful resource for the American English market, consider what this means for languages like Kituba, Stellingwerfs or Seeltersk. Consider what it means when you want to use a translation dictionary for such languages.. These resources become a reality when there is the necessary cooperation; this is what OmegaWiki hopes to achieve.

There is a need for people to work on their terminology and indeed we would like to include the terminology of falconry, tennis, and ships. It will happen when it does.

Q&A
*OmegaWiki does want to include out of copyright content as well. For us it is a start. Collaborating with for instance WordNet would be however more important and relevant.
*Proper names; yes we want them; we have George W. Bush already for quite some time.
*We would retire words by indicating them with a date indicating when they went out of use
*Circular definitions are even more problematic in OmegaWiki, this is a great example why.
*Context .. yes, I wish this was a problem that we have to deal with.. we need more functionality
*Print versions .. this is at this stage no issue. Nobody has indicated that they want to work on this.

Conclusion:
I really enjoyed Erin's presentation. It helps me to get my mind around issues that have not popped up for OmegaWiki. As many are quite will make their appearance, it is best to be forewarned, it allows us to get forearmed.

Thanks,
GerardM

Thursday, February 01, 2007

Google defuse the googlebomb .. GREAT

When you are of the opinion that George W. Bush or "Shrubya" is a "miserable failure", you could find confirmation for this by googling for this and this truth would be on top both in the Google, Yahoo or Microsoft search engine. Effectively it is a prank. It is not what you really want to find and Google announced it has worked on an algorithm that will prevent a Googlebomb in future. Effectively making the word a misnomer it is now more correct to call it a Yahoobomb or even better, a Microsoftbomb.

In an article in the Guardian, the fear is expressed that by manipulating rankings in this way, Google will exert its power and be able to manipulate what is seen as true. At issue is that the reason why Google defused this bomb was because people believed it to be true because Google said so... (there is no such thing as common sense as common sense ain't common)

In many content projects, bots create links to websites they hope to make more relevant in the eyes of the search engines. This type of vandalism resulted in a backlash where Wikipedia now indicates to search engines to disregard any and all links and thereby invalidating the basis on which search engines operate. When Google were to have algorithms that filter and punish this type of SPAM, it would lead to a more sane environment.

Google is open about its intentions. Microsoft is open about its intentions as well; as long as it discriminates against its competitors like Wikipedia I will be happy to use Google knowing that Google is kept honest by having competitors.

Thanks,
GerardM