Tuesday, February 28, 2006

What would you have done that makes sense

I talked with SJ the other day. We talked about many things but the things that come back to me is what would we have people do if they want to spend serious effort on things that make sense to the aim of the Wikimedia Foundation.

The aim of the Wikimedia Foundation is to bring all knowledge to people in their own language. The aim is breathtaking.. It is absolutely audacious, how do you go about making this happen. In a way it is a journey you embark upon and there are many small things along the way.

So what are the things that make a difference, things that can be done within half a year. How about creating fonts, fonts for languages that do not have a Free font yet. Or even define the script for a language that does not have a script. How about creating inflection boxes for parts of speech for WiktionaryZ? How about thinking wildly how you could do something you take for granted and do them in a new way. How about writing documentation for MediaWiki of for a project (Wikipedia; das Buch I do recommend :) ).

How about writing software to ease the translation of Wiki content? This could be by having OmegaT read and write directly to a MediaWiki resource. How about having people work on content that is underdeveloped. Yes, the English Wikipedia will have a million articles, but where is the content in Swahili, Farsi, Hopi ? Even a million articles will not tell you about all the villages in Ghana, Honduras or Belarus.

These are just some of the things that I can come up with without trying. What would you have a student or a few students do when they have half a year to work on a "term-project" ?

Thanks,
GerardM

Thursday, February 23, 2006

Upper or lower case in Wiktionary

Yesterday Brion made an executive decision. He decided that ALL wiktionaries will have lower case articles from now on. He changed the database so that lower case articles are allowed and ran a program that changed all entries from uppercase to lowercase.

The result was predictable, some wiktionaries like the Norwegian is really happy. They were anticipating this change of their database. They were ready for it and the major problems were done in a few hours. Not so with some of the other wiktionaries, there this was completely unanticipated and a lot of feathers were ruffled. When you do not have a plan, when it is against what you wish it does not make for an effective change.

The other thing that was not really pleasant, is that the bot used to create the interwiki links is broken. It is broken yet again. It does work after a fashion but it does not do all the work that it is supposed to do.

This month is probably set another record for the RobotGMwikt bot. As so many wiktionary entries have changed, it has to be changed on ALL projects. This month may be good for some 500.000 edits.. It is indeed great that we have interwiki links but the technology is not efficient.

Thanks,
GerardM

Saturday, February 18, 2006

White spelling

When you consider Wikimedia projects, one of the leading guidelines is the NPOV. The Neutral Point Of View is one of the most important things in dealing with differences of opinion. Particularly for Wikipedia it is really relevant. It gets the sting out of many arguments because what you can quarrel over is limited. Prove your point and show that it is a fair representation.

Getting into an argument about NPOV on a Wiktionary is different because you define a concept well and when someone is of the opinion it means that you define another concept. Technically they are DefinedMeaning in WiktionaryZ. So today I had one of these NPOV situations. Something to do with my mother tongue.

The official spelling is published in what is known as the "groene boekje". The latest version was published in 2005. There are a few problems with it; it is a proprietary list so organizations like Open Office cannot get it to build new spell checkers, it is also not available as a list of Expressions for WiktionaryZ. The datadesign is such that it allows for spellings that are correct according to an authority.

The other big problem is acceptance. Yes, it is the official spelling, but what if people and certainly big publishers do not accept it ? There is a new movement called "Witte spelling" that intends to create an alternative that is less confusing. This will result in a list of words spelled correctly according to this list. It results in a "green" and a "white" spelling. When we get the witte spelling as a resource, we can create a spell checker for Open Office, we can inform about the correct spelling according to the while spelling..

From a NPOV point of view, doing it in this way is problematic. The official spelling gets underrepresented, but how can we do it justice as it is proprietary? In several way the official spelling becomes less relevant..

If anything this is a great example that making what is supposed to be a standard proprietary, is a self defeating strategy.

Thanks,
GerardM

Thursday, February 16, 2006

Building a community of developers

The number of people who work on the MediaWiki software is limited. There is a growing need for functionality. This need is the consequence of MediaWiki being considered the best of breed of the Wiki engines. It is a consequence of the use of wikis for knowledge management, it is used for documentation efforts.. MediaWiki is used a lot.

Many needs for improvement arise from within the Wikimedia projects. These needs are typically taken care of by people who scratch their own itch. I am particularly interested in stuff when it helps me with the WiktionaryZ project. Other people have a need for Wikinews or WikiSpecies functionality.

One problem with the current model is that the developers are nominally all volunteers. This is when you analyze it no longer working it is also not true; the best developers are being snapped up by organizations and consequently are working either for interested parties or they are no longer available for MediaWiki work.

This means that it becomes more and more difficult to get things programmed. WiktionaryZ is an ambitious project. It needs available programmers and it needs people who can program and know other languages than just any "European" language and English. This need is felt more and more acutely.

I hope to develop contacts that I have made in Africa. A programmer that came recommended to me by someone who manages a great project in Swahili as best as he can. Erik did write this nice specification for something called InstantCommons. We hope them to develop this for us. When this works out, we have some new MediaWiki developers.

As we typically do, we discussed what to do when this is a success and, when we have a need for MORE developers.. Because of my interest in Iran (Farsi and Luri), I said that trying a similar would be a good idea. This got me in an interesting argument; Iran is with its present policies and president seen more and more as an enemy. Some people consider this to be a reason not to use the contacts that exists. From my perspective, this is not rocket science and it is about words and how they are used and understood. If anything we should collaborate on this.

Another strategy that we could adopt is having a competent developer, someone who is also a good communicator help students when they do termprojects related to MediaWiki. Any project related to MediaWiki. I think up to 50 projects could be handled given that a project takes half a year.. This would mean on average 100 students.. The benefit would exist in two ways; a lot of work can be done in this way and we probably would have a retention rate of something like 5% of the people as developers for a period of minimally a year.

Building a community of developers is essential. It is however not that easy :)

Thanks,
GerardM

Friday, February 10, 2006

What is a language

When I was in school, I had to learn several definitions for "intelligence". My favorite still is: "Intelligence is what the intelligence test measures". Many people use as their definition for what a language is; "a language is a language when it has its language code".

From a technical point of view, such a definition is beneficial. For all the major languages it is simple; it is obvious that there is a language code that describes them. For some languages there is a language code because people in the west take an interest, tlh is one code for one such language. For other languages it is more problematic, some of them have their code but that can make the problem worse. Some tools rely on the codes to be their and applicable; OmegaT, a CAT tool, for instance relies on the codes that exist in its programming environment. This is a serious problem because this programming environment supports ISO-639-1 and only with ISO-639-2 the code for the Neapolitan language became available. Consequently a translation tool does not support MANY languages. Even ISO-639-2 does not really help; the Kurdish language is acknowledged not to be a language; it is considered a language family that consists at least of three languages. These languages are acknowledged in the ISO/DIS-639-3.

While the ISO/DIS-639-3 is a huge improvement, it gets opposition from a few quarters. Many people, particularly developers of software, are of the opinion that some 8.000 languages is too much. Other people are of the opinion that the number of languages is not big enough. But also some languages that had support can be considered a dialect of another language like what is happening for the Twi language. How this will be appreciated by the people who speak Twi is anybodies guess. Twi is considered to be part of the Akan language, this article on the Akan language is indeed another example of the systematic lack of attention Africa gets.

For a CAT tool, it is really relevant that it allows its users to use the tool to its fullest potential. This does not mean that standards should not be supported, it means that multiple standards should be supported AND that you can introduce user defined languages as well.

Thanks,
GerardM

Wednesday, February 08, 2006

Gmail now with instance message functionality

The great thing of Google tools is that they are useful. Not only useful but they have a knack of taking something that everybody does and then add something to it that makes it better. Gmail was great; all the mail that you receive online AND presented in a way that makes sense.

When you receive as much mail as I do, most of it overwhelmingly from the same "source", mailinglists. Many people are subscribed to the same mailinglists and the reason why most people use Gmail... It would be cool if it were possible to identify mail addresses like the WMF mailing lists the point is that they can be stored differently. The content is not personal and it would be nice if it could be treated differently, it would be nice if mailing lists could be identified and stored separately.

The great news today is the new chat functionality that was added today. People who do not use skype or IRC but who have Gmail now can be chatted with. Two people I have communicated with for quite some time, now are available for a chat.. really powerful and guess what, one of them uses Google talk.. That was really sweet.

Thanks,
GerardM

Wednesday, February 01, 2006

MediaWiki is secure software

MediaWiki, the software behind projects like Wikipedia and Wiktionary is software written with an eye for security. The practices that are employed prove that security is taken seriously. The point is that people do not always understand what security means and what security is provided.

Many of the MediaWiki implementations allow everybody to create and change articles. This is a conscious decision, it is part of the formula and consequently this is not a problem from a security point of view. As a consequence the problem of maintaining quality content and preventing people vandalising the content, is a management problem. The tools to manage this problem are diverse but many tools that are considered security tools are usable.

Often vandals do not know that what they do is useless. Often people add links to all kinds of websites in order to increase their Google-rating. The MediaWiki software indicates to the Google crawler NOT to include external websites for its ratings.

Blocking IP-ranges and users because of persistent vandalism is one. Trusting logged in users more than anonymous users is another. There are many Wikimedia projects and all of them still have at present their own users. In Februari, it is planned to develop single signon for the Wikimedia projects. Single login has been on the wishlist of many of the people who are active on multiple projects.

With single login, in essence a management issue with security implications, it becomes feasible to use this as a stepping stone for the implimentation of security features that help with the management of vandalism.

The feature that I would love best is to differentiate the strength of authentication based on where a user comes from. When a user comes from a school with a history of vandalism, it makes sense not to allow anonymous edits. There are many of these types of soft security measures possible.

On mailing lists about Wikimedia, there was talk about a patch that allows for logging in users who authenticate themselves with OpenID. The interesting thing was that people had two issues with this; first it would not allow me to use my MediaWiki ID as an OpenID. The second is that to some extend OpenID is going to fit into the YADIS framework (Yet Another Decentralized Identity Interoperability System).

Yadis is interesting because it is linked to the eXtensible Resource Identifier or XRI, a standard that is developped by OASIS. It is also linked to the W3C (YASB - yet another standard body :) ).

In the end it comes back to standards; when the WMF would support twoway YADIS authentication, it makes for a VERY relevant implementation of security related functionality. This could provide for better management in the fight against vandalism. It is however important that it is a standard that we provide. That is why I am of the opinion that the WMF should support standards.

Thanks,
GerardM

Tuesday, January 31, 2006

Rule number one: You are wrong !

There are those days when you must be wrong .. Today was one such day: because the vision that is behind WiktionaryZ "appears to be based on a simplistic and Eurocentric view of language". It must be true because that is what is being said .. If not, rule number one must be applicable ..

Today a nice standard called OLIF was pointed out to me. OLIF stands for Open Lexicon Interchange Format, it is an open standard for lexical/terminological data encoding. It is essentially a Free and Open standard, it has many illustrious and industrious backers. But it is a tad Eurocentric. It now has East Asian language support, but it is best at English, German, French, Spanish, Danish and Portuguese.

I would not want to dismiss a standard like OLIF when people are actively involved in a standard .. To be useful would be necessary to define how a standard is lacking, it might be that it is just a matter of some refining. It might be that the standard does consider things that I have not considered yet (and I do know that I do not consider everything all the time).

On Meta we are voting for a new Wikimedia project, this project is about standards .. Discussing standards but more importantly it is a conduit for working on those standards that do not fit the requirements that we have for standards in the Wikimedia Foundation's projects. Now there are two options: I am wrong or the standards will fit our needs like a glove.

Thanks,
GerardM

Saturday, January 28, 2006

About new words

More and more I find how my appreciation of things to do with words is changing. By chance I found this book Taal van het jaar vijf It is a publication by van Dale, the Dutch producer of quality dictionaries, it is a collection of 52 articles describing new words that make their appearance in the Dutch language. With my interest in words and more I find it a superb collection of observations how a language changes by the inventiveness of the people who use it..

If you know Dutch, you will love it. If there is more of this, also on other languages, on the Internet .. I would appreciate to learn about it :)

Thanks,
GerardM

Monday, January 23, 2006

Standing on the shoulders of giants

At a conference in December I went to, I received promotion CD for the "Referentiebestand Nederlands" of the TST-Centrale. Yesterday I spend a considerable amount of time digesting the information on the CD. I learned a lot, particularly the power of having content in a database that can be used in many different ways. The "referentiebestand" is a corpus based dataset with 45.000 Expressions with morfological, syntactic collocational, semantic and pragmatic descriptions. (much of the definitions are considered stubs in the English Wikipedia, please help in describing these subjects better).

I have learned a few things. I am right when I understand that the location of words in a sentence is relevant. The TST-Centrale uses a fixed notation for it; I wonder how universal it is.. On the CD it is explained that people who use this content, can select the information they want to use and that this is part of the secret of its success. We hope that having the WiktionaryZ available as a database will serve in a similar way.

Reading back the first paragraph, I feel like a Tom Thumb. Every second noun is something to look up and it was hard work writing it. The great thing is, that when you have it in front of you, when you see it demonstrated it does make sense. When we collaborate on content and stucture, when we make WiktionaryZ something that is usefull because it has an application, we will have giants that help us little people make progress when we are allowed to stand on their shoulders and move with us forwards.

Thanks,
GerardM

Saturday, January 21, 2006

A discussion on trusted computing

Today I had a big discussion about trusted computing. The question was is it bad is it good. Should we be against it and why.

From my perspective the biggest problem with trusted computing that I have is that it may be an open standard, but it is not a free standard. The specifications of the standard that the trusted computing group is available for organisations, it costs at least $1000,- and thereby excludes all these people that create wonderfull programs that people share. Because of this lack of openness it is has a fundamental problem. It makes me trust organisations that I not necessarily trust; why should I trust Yahoo, MSN and AOL as they demonstrate what I perceive as a lack of protection for the privacy of their clients? This trusted computing architecture does not allow me to trust my own software of the software created by a friend. It does not because I do not see how I can create software that will be trusted by my system.

Personally I do not think trusted computing is the equivalent of digital rights management. I am of the opinion that DRM leads to giving away rights that are mine.

Trusted computing does one other thing. I expect that it will take away much of the anonimity that is still with us on the Internet. This aspect is probably something that few people considered. My first clue came after I realised that it is a perfect tool to do vandal fighting on Wikipedia. It gives us a tool to more precisely know where these people come from and it can give us much better protection against this scurge. The other side of the coin is that when it provides us with more control, it will also give more control to those that I do not trust to use it wisely.

Thanks,
GerardM

Sweet nuggets of relevancy

The Kamusi project is one of those deserving projects, it is a project dedicated to the Swahili language. It intends to create a dictionary. A dictionary that has been worked on for many years. A resource to be proud of.

It is only lacking in this one resource.. money.. If it does not find money it will pack it in. They are now finishing off software so that the project will be in the best shape it can be. I truly hope that the Kamusi project will find its way into a bright future. If you can help, please do :)

Thanks,
GerardM

Friday, January 20, 2006

etymonline.com

Giving good content a home is expensive. When you have good quality content, the people will come. Two truisms that are obvious. Dvortygirl mentioned the etymonline website today. In this website you see both truisms in action. It is a great resource, people love it and there is a problem hosting the great content because of the cost associated with a great user experience.

As it says on the information on the mainpage, they only have some 20.000 visitors, the size of the database is "only" 51,08 MB. My reaction to these problems are predictable; I would like to host this information in WiktionaryZ when we are ready for etymological content. There are however a few issues. These issues have to do with the perception of Wikipedia.. Let me quote:
  • "Approach Wikipedia about a partnership, or actually merge the site into Wikipedia. This is a painful option. In a sense, this site is the anti-Wikipedia. It is deliberately not open source. You'd understand why if you saw the regular stream of e-mails I get from people insisting that their own crackpot folk-etymology idea is absolutely correct. Such things can be based on some fervent politico-racial agenda, simple insanity, or "my French teacher in 10th grade told me."

    A major reason this site exists is to serve as a template against which to measure people's best guesses and wacko theories. The whole Internet is a big Wikipedia; this site is a compilation of the most rigorous academic information."

I would have several answers but the most relevant thing is Wikipedia=Crackpots. Yes, our project is not Wikipedia. Yes, we want well researched information. No, it is not Open Source it is Open Content. All this does not address the issue of the quality of content.

In WiktionaryZ we have a need to address quality. In the current Wiktionary we have our crackpots. There is this loon who insists of there being a word called "exicornt". He even threatened admins in order to have this word exist in the Wiktionary.. :( We have to deal with these persons.

We could do something for etymonline. We can approach them. We can offer to host its content by importing their content (obviously with proper attribution), but we sure have to address the issues. Quality is important and we have to protect quality content for our own sake. When we do, and when are successful opportunities like this will be less problematic.

Thanks,
GerardM

Thursday, January 19, 2006

NPOV and sources (continued)

In the blog entry titled NPOV and sources, I mentioned the need for one resource where the sources are given for the Wikipedia articles on one subject. The consequence is that you may end up with a resource that is truly big. Today I learned about a resource of the UNESCO that has 1.600.000 references in the Index Translationum. The Index Translationum is a resource of books and its translations.

This is marvelous resource; there are some caveats. The on-line data is only from 1979 onwards while the resource started in 1932. The printed edition is available in the UNESCO library in Paris...

There are many possibilities with a resource like this.. It is indeed one of the resources that would benefit on a massive scale from a GOOGLE Book action. Obviously I would not care who does the scanning and who spends effort on digitizing this conten. It is however an extremely important resource and everyone would benefit if the data from 1932 to 1979 becomes publicly available as is the intention of this project.

There is more to learn about this project. When considering using it for presenting our sources in Wikimedia projects, we would like an API to refer to it.. I have not looked into it.

Thanks,
GerardM

Wednesday, January 18, 2006

What do you do when your computer is broken

Well, you ask if you can use someone else´s computer. I did. So then I used my mother´s computer.

What do you do when the WIFI router does not work anymore .. You use the ethernet option this router provides...

What do you do when the keyboard mapping is broken .. you do not use question marks ..

Life is a bitch.

Thanks,
GerardM

Monday, January 16, 2006

What to write about ..

With the new WiktionaryZ blog, topics about WiktionaryZ will go there. There are many topics that I have been writing about that are not directly related to WiktionaryZ. These are topics like IEEE-LOM, fundraising, Wikidata, how I think we can improve Wikimedia projects..

Advertising

Yesterday Amgine wrote an “Advertising proposal. The idea is that we can advertise the services that the Wikimedia Foundation provides. Essential for good advertising and marketing is that you target your audience. With such an attitude comes the focus that would improve our content. The advertising would be done using an advertising server; the many people that have their own wiki, their own blog, could include advertisements from this server.

Wikidata

The Functionality that was originally conceived for Wikidata will end up in Mediawiki it self; this is both a blessing and a curse. The great thing is that it does signal that much of the functionality that we conceived is indeed relevant and it will bring functionality to all the wikis that use Mediawiki. The drawback is that it has implications for the design. It does complicate things at this stage, but on the other hand when we have a great instead of a good design from the start, the extra time and effort needed now may pay itself back in the long run.

Thanks,
GerardM

Sunday, January 15, 2006

Our blog

Words and what not is my blog. Now that we have a "commission", my blog will become different. It must become different because it will be important that it is understood that all these people have their own opinions, ideas and that this commission is also a process. A process of getting this thing on the road, a process of developing a consensus of what we are doing.

All these things are not new. It happened all the time.. Now we will make this process more visible; it is done to read this other blog, the WiktionaryZ.blogger.com blog of the commission.

The thing I am not sure about is to what extend I will continue blogging here and to what extend I will blog on the new blog.

Thanks,
GerardM

Saturday, January 14, 2006

Presenting: the Commission

WiktionaryZ is a big project. It is becoming bigger. It is an innovation on the Wiktionaries, as such it is a continuation of these projects. It requires innovation in the MediaWiki code, the first software is about to be merged into and become available in all MediaWiki wikis. There is the first alpha version of the WiktionaryZ software, and it hosts the GEMET data. We are talking with several relevant organizations that share the dream that still is WiktionaryZ.

There is one issue; people think that the problem with this project is that it is run by one person; me. It was once put to me was that that I represented a "truck factor" for the project. From my perspective, WiktionaryZ is certainly my dream, but it has never been my dream alone. I have worked hard to make WiktionaryZ happen, but I am not the only one that worked hard to make it happen this far. I came up with many ideas that made WiktionaryZ what it is, but not all ideas have been mine, all the ideas became WiktionaryZ because of the many conversations, e-mails and IRC chats about them.

From my perspective, there is this issue that I do not scale. There are things that amount to policy and it needs to be expressed because policy dictates technological choices and we are building the technology for WiktionaryZ. This implies that some policies have been set but it also implies that more questions will raise their head that do require an answer and require an answer quickly because we are building the software, the network, the connections now.

I did discuss this issue with Jimmy Wales, and he came up with the great idea to have a commission for projects like Wiktionary. This commission could have several funtions; it can act in a similar way as the chapters do for countries for projects. One role of the commission would be that the policies of a project would be consistent with the aims, the policies of the Wikimedia Foundation. Another equally important part would be to represent the community that will make WiktionaryZ its own. To do all this, members of such a commission have to be part of the discussions about the developing Wikidata and WiktionaryZ.

I like the idea and, I have asked several people to become part of an initial commission. They are all Wiktionarians (with one exception), they represent many Wiktionary projects and they are and have to be communicative; they do use Skype/VOIP and often can be found on IRC.

The commission:
Including Sabine Cretella is obvious. Sabine developed WiktionaryZ with me from the start. WiktionaryZ is a dream she has fostered for a long long time. Sabine is active on the Italian Wiktionary and is one of the initiators of the Neapolitan Wikipedia. Sabine is a professional translator and is known in this world as an evangelist of Open Source and Open Content.

Of all the Wiktionaries, English is the most relevant. Dvortygirl has been active there for a long time. Like me, she is an admin and she is well liked and respected for her work.

Gangleri is active on many wiktionaries. The thing I really appreciate is his involvement with right to left languages like Yiddish. On the Internet, these languages have their own issues, Gangleri is active in the Mozilla organisations to address several of these.

Yann, is active on the French, Hindi and Gujarati Wiktionary. He is also the treasurer of the French chapter. One of Yann's challenges is to help us get more people interested in the languages from India.

Erik
is the one who is not into Wiktionary. The reason why he is invited is because he is the architect and realiser of Wikidata. This is the enabling technology for WiktionaryZ. Erik is also the realiser for WiktionaryZ. Wikidata has in WiktionaryZ its first application. As this is a truly big and complex project, many of the things that would hit Wikidata eventually, need to be addressed from the start. Erik is also important as a linking pin to the Mediawiki developers.

GerardM, if I need an introduction I invite you to read this blog.

Some policies or, a glossary of our policies:
Availability: "our data is to be made available through open standards and in a non-discriminatory manner"
Data design: "WiktionaryZ is implemented as a relational database. When information that is relevant in the context of WiktionaryZ cannot be added, we will try to ammend the design."
Full functionality: "WiktionaryZ needs to be able to include the information that is available in the Wiktionaries"
Partner: "a partner is an organisation that collaborates with us in the realisation of what we intend with WiktionaryZ"
Sponsor: "a person or organisation that donates money or content to the project or to the Foundation"
Success: "success is when people find an application for the WiktionaryZ data that we did not think off."
User Interface: "the interface should be in any language. We want this both for the Mediawiki and the WiktionaryZ user interface"

Thursday, January 12, 2006

Dialects within a language

There are Wikipedias in many languages. So far there are some 212. Most of these languages have an ISO-639 code. There are two versions of this code that are "official", there is one version that is workable; the ISO/DIS-639-3 is currently maintained by SIL international. However, workable does not mean that it is perfect. Today there were to moments where the current practices relating to ISO-639 were the issue.

JAVA uses ISO-639 for its language codes. The codes used is the ISO-639-1. Consequently the Neopolitan language is not known. OmegaT is an open source CAT tool, it uses the languages known to JAVA as the languages that it can translate.. So in order to translate to Neapolitan you have to pretend that it is a different language.. Not nice.. So the nice people of SUN were asked this and we have great expectations.

Today there was a request on Meta, the website about the Wikimedia Foundation's project, for a new Wikipedia. The request is for tarantino, it is considered a dialect, a dialect of Neapolitan. This request is problematic because there is not even an ISO-639 code. Consequently there is little chance of there being a wikipedia for created. Now, with the new namespace manager, it is possible to create a seperate namespace within the nap.wikipedia.org for the tarantino dialect. This is also a solution for the problematic request for a Lower Saxon wikipedia that will be in an orthography that is not German..

It is sobering to see that standards can enable and prevent things to happen. Good standards are vital and ISO/DIS 639-3 is a big move forward.

Thanks,
GerardM

Wednesday, January 11, 2006

Machine translation

Machine translation is a difficult thing. It is hard to get it right, the English language Wikipedia has a problem with writing a good article about it. Some people think that with WiktionaryZ we are in the business of machine translation, as a result the subject gets on our radar screen every now and then. It would be a good characterization to say that the Machine Translation we could use is the one that is able to translate the definitions of a DefinedMeaning in all languages. This resulting translation should be good enough so that people have an idea what is meant .. in Bambara, or Papiamento... Translations like these are what we need and there are as far as I am aware no Machine Translation engines.

To create a Machine Translation engine for these language, you would need all kinds of rules about the languages that you want to translate to. Now I would like to know these rules. Not so much to build Machine Translation engines but because they have this other application, one that is much closer to my heart, it is needed for software that teaches people languages. The Universität Bamberg is working on exactly this. They want to use the data of WiktionaryZ for this purpose. So if you have a nice sets of rules for them I would be obliged.

Thanks,
GerardM