Wednesday, September 06, 2006
Lies, damn lies and statistics
Many statistics seem to proof what the researcher tries to prove. Aaron Swartz wrote a really nice article on who writes Wikipedia. It is great because it challenges conventional wisdom. The conventional wisdom is that a small group of people write Wikipedia. Aaron makes it plausible that it is the anonymous user who contributes the most letters to the article that you read in Wikipedia.
What is particularly important for me are the consequences for WiktionaryZ. The suggestion is that we have to make it easy for the casual user. That people contribute to the things they know and care about. That an intuitive screen helps, WYSIWYG helps people who are the writers, the wikisyntax is for editors, the people who make things pretty.
The user interface of WiktionaryZ needs to be compared against the user interface of Wiktionary. Wiktionary is flat file and almost every Wiktionary is essentially different. WiktionaryZ will have one user interface for all languages. Contributions made by some will be available to all. I trust that we have and edge.
The flipside of the coin is that indeed a limited group of people do the EDITING. What we do not have in place are tools that help editors. Editors are the people that will prevent WiktionaryZ from becoming a mess. Particular the merging and deletion of the [[DefinedMeaning]] and [[Expression]] will be important to get right..
There are ideas on how to do this, they are not mature yet.
Thanks,
GerardM
Monday, September 04, 2006
Colours
Lemon yellow, is such a colour. It has the RAL-1012 code. Because of it's long existence, the colour indicated by this code has been translated to numerous languages; suurlemoengeel is the name of the colour in Afrikaans. When the colour is defined by the RAL code, and you use the names used with this colour, you have an identical meaning for the word. This is obvious as this is what the RAL codes are there for.
The question is, how to list the RAL codes themselves in WiktionaryZ. The RAL colours are a collection, we can describe them as such. We can also include the RAL-codes as an expression, the zxx language seems obvious to me. We can do either and we can do both.
Thanks,
GerardM
Sunday, September 03, 2006
The Chinese language
The basis for this is the way we have embraced ISO-639-3 and the experience we have gained with Serbian and English. For Serbian we have two scripts and this works well, for English the words that are universal are English, the specific US-American words are now English (American).
The only thing left for English is to include English (British).
Thanks,
GerardM
Sunday, August 20, 2006
Latin and species
It was suggested to me to have Latin. The motivation was; I have this list of birds and would it not make sense.. It does make sense on one level and, on another it does not. The Latin used to come up with all these taxonomical names is not necessarily the kind of Latin that helps you learn the language.
When we want to include THAT kind of Latin, it makes sense to include the taxonomical relations and attributes that makes taxonomy the science that it is. Considering this, it will need Wikiauthors for it's publication data. It will need some specific database functionality to make this feasible..
I do want Latin but I doubt if this is a good moment to include it.
Thanks,
GerardM
Friday, August 11, 2006
After Wikimania ... semantic mediawiki
When I was in Rome I met Denny. He demonstrated how well Semantic MediaWiki can be used for personal use. I was really impressed. Denny installed all the necessary bits on my computer. And it does do many of the things that I need. It allows me to create multiple lists associated with a subject. For instance the organisations involved and the persons involved with a project..
The more I see it, the more impressed I am. The only potential problem I see is that people have to learn more wiki-syntax. As it is so powerful, I would not mind to include semantic mediawiki in WiktionaryZ.
Wednesday, August 09, 2006
Yochai Benkler: The Wealth of Networks
- The message was heard for a first time by a public.
- The message was told for the first time.
- The discussion following a presentation brought new insights.
- The presentation raises questions.
As I did not grasp what the new role for organizations will be in this brave new world, I bought the book. I bought the book because I strongly believe that there will be a role to play for organizations and if there is to be a recipe for including organizations, businesses I want to have it. If there is no such recipe ..
I have browsed the book so far, many things I do not grasp. In discussions I had earlier in the week I learned how different the notion of "liberal" is depending on the context. As this is also a strong theme in the book, I will be struggling. However, it is fun.
Thanks,
GerardM
Sunday, July 30, 2006
WiktionaryZ and attributes
The good thing is that it is an indicator of things that may one day make our day. The thing however is, that it does not necessarily have our highest priority. At this moment our priority is to get our versioning working. We need to have historic data, we need to have improved information in our recent changes.. This is what we need badly.
We need it badly because this is the core functionality of WiktionaryZ or Wikidata. It needs to be done. It needs to be done well. It needs to be done before we add even more complicated functionality in WiktionaryZ.
There is a lot of work that needs to be done. We want to have proper support for terminology, we need to include the notion of "domains" for that. We want to have proper support for lexicology, we need to include the notion of language dependent attributes for that. We need to have relationtypes that can only be chosen in the correct context. The context being either the language or the domain.
All these things will happen. They will happen in a way that is consistent with the amount of resources available. This is why we have a need for involvement. This involvement can come in many different ways from from many different places providing many different angles to improve our data. The key thing is that we will maintain our architectural integrity. It is completely unacceptable to build all kinds of wished for functionality while the technical foundation has not been laid.
The collaboration that became WiktionaryZ will start it's third year at the end of August. There were many reasons why it has taken this long. What it brought us is a design and a philosophy that may work. The third year will be the year where I expect that we will have our full functionality.. When we do, things will evolve more quickly.
Thanks,
GerardM
Tuesday, July 25, 2006
The operational definition meets the DefinedMeaning
When such an operational meaning cannot be correctly identified by 90%, it means that all the definitions of the different concepts are suspect. It may mean that too many concepts have been identified; certainly more work needs to be done to define things better.
In WiktionaryZ, many DefinedMeanings may exist where it has been indicated that the definition does not define the concept really well. These concepts are close but no cigar; they should not be taken into account when it is determined if the Definitions are indeed operational definitions.
The question is, how do you then identify the quality of the translations and the usability of the synonyms.
Thanks,
GerardM
Sunday, July 23, 2006
Being too busy
The thing that amazes me is that I am so busy making sure that everything works well, that we have some crucial data. Things like the Swadesh lists and the the list of 1000 basic English words are really important because these are basic words, they demonstrate best the merits of the concept. And I am really pleased with how things are progressing.
The thing that annoys me is that there is so much work that I would like to get done.. Some of it is plainly not for me to do. Other stuff like writing on the blog is very much for me to do. I may have to be even more selective what I do. Making these choices is hard.. Then I am glad I am in this position.. So many great things are happening and I am part of it :)
Thanks,
GerardM
Tuesday, July 11, 2006
Water; at least three DefinedMeanings
It has been said before by many, it is the simple words where "everyone" understands what is ment which are the hardest to define.
Thanks,
GerardM
Saturday, July 08, 2006
Just a word
In Dutch the word has a clear agricultural background. It refers to how a weed, once it is established seeds itself to neighboring areas. It is one of those word that you hate to hear in another domain. Today I did. There is so little that I can do, I feel sad and can only hope for the best for Scott.
Thanks,
Gerard
Sunday, July 02, 2006
A dubious record
The bot is quite nice, it can run autonomously and it updates some 30 wiktionaries at the same time. For the smaller wiktionaries, I run it specifically for a project every now and then.
Yesterday I chalked up the 200.000th edit for the English Wiktionary. It makes it nor unrealistic to think that I have some 600.000 edits on all the wiktionaries. There are six instances of the bot that run at any one time, they run on different projects. I think it is a dubious record because it does not bring me any happiness; it only indicates that there is information about a word that is written in the same way. I doubt very much that anybody does anything with it.
For your amusement; there are MANY English words on the Chinese Wiktionary that cannot yet be found in the English Wiktionary .. :)
Thanks,
GerardM
Saturday, July 01, 2006
What is your mother tongue
Now what is the mother tongue of Sabine's kids ? They speak primarily Italian...
In WiktionaryZ we have people that indicate that their mother tongue is zho or Chinese. According to Ethnologue Chinese is a macrolanguage. This implies that Chinese cannot be a mother tongue, one of the 13 languages Chinese is divided in can only be the mother tongue. This is a potential hot potato when people equate Chinese with the country and not the language.
WiktionaryZ is about languages and only about languages.
It can also be understood differently, they may mean that Chinese is the first written language that they learned. However if I understand things well, when people talk about the Chinese written language, it is actually Mandarin. For Yue for instance, there is a need for additional characters that are in one of the later versions of the UNICODE. This is however not what I would consider a mother tongue. A mother tongue is the language that you learned from your mother. Writing is what you learn at school.
Thanks,
GerardM
Saturday, June 10, 2006
Punjabi and what IS that script
I then tried my hand at the Punjabi language. Punjabi is written in two scripts, and I guessed wrong trying to identify them. One was indeed an Indic script, ਪੰਜਾਬੀ is written in the Gurmukhī script while پنجابی is written in the Shahmukhi script. Shahmukhi is indeed an Arab script but it is not the Arab script. In order to properly identify these words, I looked them up at Unicode where there is a nice list of the ISO-15924 script codes. Gurmuki has Guru as its code and Shahmukhi .. is absent.
It is probably pretty safe to indicate it as Arab, but when my information says that it is not, it is indeed problematic. I could also have it as an uncoded script. The problem with these standards is that they work up to a point. The point is what are they there to do.
When you write Dutch the standard Latin script is used, however that leaves out one character and consequently all word processors capitalise the ij wrong, it should be IJ and not Ij.. I think a similar thing is happening with the Shamukhi script. It is assumed to be Arabic but the style of the glyphs is different. I think it is just one of those things that may change in the future.
I think I will indicate it to be Arab for now .. :)
Thanks,
GerardM
Sunday, June 04, 2006
What to do next
The question is what to do next. We cannot and should not rush the development but we do have a host of data that we could import. It could be other thesauri, glossaries or ontologies. It could be a long list of Expressions in a given language to populate a spell checker. We could import the data resulting from Duesentrieb's Wikiword application, this could give us a link to Wikipedia articles.
When we had more active collaborating developers, we could consider doing the import and export routine or we could start working on inflections.
What would you consider the next bit of data to import after the languages? What would you start programming on given where we are at this moment?
Thanks,
GerardM
Wednesday, May 31, 2006
What use is a community for others ..
One presentation, the price winners presentation at the end of the conference mentioned a nice scheme were two concepts were compared and the question was, to what extend is the first concept associated with a second. People doing this are trained with a first set of concepts, they are then asked to do a further set of concepts to see to what extend they have learned things and then .. They are off. This worked really well but, you need a large group of volunteers or you use a computer. The computer did either a good job or gave COMPLETELY different answers from what a human would do (these are the things to watch our for in a Turing test).
My idea is, when WiktionaryZ gets itself a large community of people interested in languages, it would also be natural to ask this community if they are interested in helping out with research. One strategy would be to have just one person check a machine derived result, when there is a discrepancy, have some more people look at it ..
Another interesting experiment would be to test the difference between the different groups of users of English including people who use English as a second language.. What do you think ???
Thanks,
GerardM
Saturday, May 20, 2006
Languages, dialects and orthographies
One of the things that makes sense to be aware of, is what orthography a text, phrase or word is in. It is definitely something that is in a class of its own and it matters when text is to be understood in an automated way. Languages do change over time and, the recognized correct orthography changes to reflect this. The German and Dutch language both have had their fair chair of changes. The functional design of WiktionaryZ has always had a place to indicate that a given spelling is dated. The way we will export the WiktionaryZ data will be by using standards like TBX, LMF maybe RDS, SKOS or something different but standard. The problem is; how do we indicate that a given word needs to be spelled different since a given date ?
Thanks,
GerardM
Friday, May 05, 2006
Languagecodes on the Wikimedia Foundation
In the past we started using the ISO-639-1 codes for indicating languages. This list was extended with the ISO-639-2 list because it was too limited. Even so there were issues with the list and they were augmented with codes from Ethnologue. This was not often considered not enough so we created our own codes.
Now there are the ISO-639-3 codes. They have a provisional status because there will be even further extentions of the list but it highlights one thing. Us using our "own" codes is really problematic. I give you some examples and I also indicate to you why many arguments used in the discussion on new languages are wrong.
The ksh.wikipedia is for something called Ripuarian. The ksh code is for the Kölsch language. Ripuarian is considered to be a language family. The consequence is that there IS no single Ripuarian orthography, language or culture
The als.wikipedia is called Allemanish. According to the ISO-639 these are four languages. The problem is that the als code is used for the main Albanian language. The code for Albanian sq (ISO-639-3 sqi) is also considered a languagefamily. The two main variants are the Albanian and the Kosovar languages. Two other members of the Albanian language family are spoken in Italy and Macedonia
Some of the most heated discussions on the request for new projects are about the status of a language; is it a language a dialect and often the arguments are of a political nature. The inclusion of languages in the ISO-639 has been political in the past. With ISO-639-3 many of these arguments have an answer with the many new language codes that have been created.
The result is that we have wikipedias like the ku the fa, the sq, wikipedia where the language is now considered a language family and where a request can be made for recognition of a language that is part of that languagefamily. There are more projects like that that I have not identified yet.
Another "nice" situation is the Low Saxon nds wikipedia. When you look at what Ethnologue has to say about the Lower Saxon language family than you get the impression to what extend there is not really something called Low Saxon in the Netherlands. The varieties of Low Saxon of the Netherlands are all there.. They are named, have there codes..
The point that I am raising is, languages are a mess. The codes for our projects are as a consequence a mess. The procedures for new projects are a mess because of the politics and the codes we have come up with in the past. What we need are better quidelines what the relation is between the ISO-639 codes. If the WMF says it uses the ISO-639 codes the codes must be in use or they must be clearly different.
Last but not least, according to the terms of use, we are not allowed to extend the codes in the way that we do.
Thanks,
GerardM
Friday, April 28, 2006
and now for some other news ..
From May the ration will be half of what the minimum amount required for each day.
They are starving and they are the lucky ones ......
Thursday, April 27, 2006
Collaboration
The open / free business model for lexicology, terminology and thesauri is .... work together on a stellar resource and make the data available to everybody. With the data available to everybody, the big question is how to achieve the best result. For an open / free resource, the best way is by providing the data in a standard way. The first standard I want WiktionaryZ to use is the "TermBase eXchange" or TBX standard. Given what we do in WiktionaryZ, LMF and SKOZ are two other great standards.
Why use standards? Simple, our definition of success is: "when people find a use for our data we did not think of". By providing the data in a standard way, it will be available in a stable way and as a result it will be more easy for people to make use of our data. It will be easier for WiktionaryZ to become a success.
Yes, I love our "competition", but I will love them to bits when they want to be as relevant as we want to be.
Thanks,
GerardM
Monday, April 17, 2006
What to do with stuff that is good but not standard
I have been thinking about this for a few days. The answer for me became obvious; it is in the database. When we have the MisSpelling table, we can have the community identify the words that should have an Umlaut. With proper logic the representation for our public will be the proper German with the umlauts.. But the first thing is to have the MisSpellings..
Thanks,
GerardM
Friday, April 14, 2006
Diana posted an answer ..
I am GerardM .. :) You can find me on the WiktionaryZ site. I am very happy to talk about collaboration. I am happy to remind everyone how we define success for the project: Success is when people find an application of our data that we did not consider in the first place..
Thanks,
GerardM
Tuesday, April 04, 2006
Some of the best things in life are free
Typically something can be called scientific when the conclusions are arrived at in a methodical way and, the method is repeatable. This is exactly what the Open Access movement wants to bring back. In order to do this they have to wrest away the restrictions that copyright has put on scientific data far too long. A lot of bad science is the result of these restrictions a lot of wastage is the result of these restrictions.
What I learned is that many superb resources are becoming available to the world as a consequence of this movement. Open Access is a rich tapestry with many threats in many fabrics of many colours.
If there was one thing disappointing it was the lack of awareness of licenses. It is a CC license.. was the answer and people applauded. Well, Creative Commons has great licenses but it is a bit like Animal Farm; all CC licenses are Free but some are more Free than others ...
Thanks,
GerardM
Saturday, March 25, 2006
A new concept; the "Regime"
Enter the Regime, a regime would be a procedure that people can submit to. It would be optional. There will be many areas where we can develop these regimes. It may be that a regime is developed or execustes with organisations that we will partner with.. One thing, to state the obvious, our data is Free so the process will be transparant and the resulting data will be Free.
It is an idea that we came up recently, we discussed it with some and now, we are interested in what you think of this.
Thanks,
GerardM
Saturday, March 11, 2006
Babel templates and WiktionaryZ
WiktionaryZ is a fully functional wiki. This means that we can add content; we can create users, we can create templates, categories. We should when it makes sense. And it does. When we start the coming test, we will start with people who understand what WiktionaryZ is about. This means that they understand the concept of the DefinedMeaning. An other factor that will help us decide who to ask, is the languages these people master.
WiktionaryZ will for now not be available to anonymous users. People can create a user. When they add Babel information to their user page, we will learn who has expertise in what language. The templates we start with have been copied from the en.wikipedia and, we hope we will get many more templates that will allow us to have five levels; the native speaker and level 1 to 4 to indicate the growing level of proficiency.
The scope of the test is limited; priority is in learning the edit process for relational data. What does work what does not. What improvement will be needed to make us ready to meet the "great unwashed". We will also be able to work on information that will be important in later phases, things like names of languages and other terminology that is likely to end up in the user interface...
Thanks,
GerardM
Tuesday, March 07, 2006
If I had money that I could freely spend ...
Interwiki links
I hate these things. They are always out of sync. Many people spend a lot of hard work on getting them right while the problem is getting bigger and it is not solved at all. It feels like a waste. When a project grows, it has articles that need to be linked. As more and more wikipedias grow, the number of articles that need updating grow rapidly. Small projects are not easily integrated. It costs a lot of resources.
With a centralised database, we could link an article to another article and by inference it would be known to all the articles that are linked to it. As all the articles are about the same subject, we could check if article names are translations. When they are, it is a basis for linking to the lexicological content of WiktionaryZ ..
Inflection boxes
When a verb, a noun and adjective changes under given rules, it makes sense to have inflection boxes. They are generated using templates on many wiktionaries, but it makes more sense to have some software that allows us to build these boxes. Software that associates inflections of one language to the inflections of other languages for the purpose of translation.
Better support for tools like OmegaT
OmegaT is a CAT-tool, it helps translators with their work. I would do two things to OmegaT, I would have it read directly from a MediaWiki wiki and when the translation is finished, write it to another wiki. I would also have its translation glossary funtction make use of WiktionaryZ..
Yes there would be a quid pro quo, when a translator adds a word to the glossary it would be fed back to WiktionaryZ.
Thanks,
GerardM
PS What would your suggestion be ?
Wednesday, March 01, 2006
Pictures in WiktionaryZ
Pictures are great. We do need them with the WiktionaryZ content. Having said this, the fun starts. We can decide to allow only for pictures that can be found in Commons. There are precedents for it. It solves the problem of what upload rules would be allowed; they would be the strict rules of Commons.
The other thing you would get into are the illustrations that are problematic for some cultures. In my country, people have as many genitalia as in any other, they look more or less the same as in any other, they are as rare as in any other and a picture would not be problematic. In many countries this cannot be done. One great example of this can be found here. Now when you read this in some 10 days time this picture may not be there anymore. It does however show you the sensitivities that need to be considered with pictures.
Thanks,
GerardM
Tuesday, February 28, 2006
What would you have done that makes sense
The aim of the Wikimedia Foundation is to bring all knowledge to people in their own language. The aim is breathtaking.. It is absolutely audacious, how do you go about making this happen. In a way it is a journey you embark upon and there are many small things along the way.
So what are the things that make a difference, things that can be done within half a year. How about creating fonts, fonts for languages that do not have a Free font yet. Or even define the script for a language that does not have a script. How about creating inflection boxes for parts of speech for WiktionaryZ? How about thinking wildly how you could do something you take for granted and do them in a new way. How about writing documentation for MediaWiki of for a project (Wikipedia; das Buch I do recommend :) ).
How about writing software to ease the translation of Wiki content? This could be by having OmegaT read and write directly to a MediaWiki resource. How about having people work on content that is underdeveloped. Yes, the English Wikipedia will have a million articles, but where is the content in Swahili, Farsi, Hopi ? Even a million articles will not tell you about all the villages in Ghana, Honduras or Belarus.
These are just some of the things that I can come up with without trying. What would you have a student or a few students do when they have half a year to work on a "term-project" ?
Thanks,
GerardM
Thursday, February 23, 2006
Upper or lower case in Wiktionary
The result was predictable, some wiktionaries like the Norwegian is really happy. They were anticipating this change of their database. They were ready for it and the major problems were done in a few hours. Not so with some of the other wiktionaries, there this was completely unanticipated and a lot of feathers were ruffled. When you do not have a plan, when it is against what you wish it does not make for an effective change.
The other thing that was not really pleasant, is that the bot used to create the interwiki links is broken. It is broken yet again. It does work after a fashion but it does not do all the work that it is supposed to do.
This month is probably set another record for the RobotGMwikt bot. As so many wiktionary entries have changed, it has to be changed on ALL projects. This month may be good for some 500.000 edits.. It is indeed great that we have interwiki links but the technology is not efficient.
Thanks,
GerardM
Saturday, February 18, 2006
White spelling
Getting into an argument about NPOV on a Wiktionary is different because you define a concept well and when someone is of the opinion it means that you define another concept. Technically they are DefinedMeaning in WiktionaryZ. So today I had one of these NPOV situations. Something to do with my mother tongue.
The official spelling is published in what is known as the "groene boekje". The latest version was published in 2005. There are a few problems with it; it is a proprietary list so organizations like Open Office cannot get it to build new spell checkers, it is also not available as a list of Expressions for WiktionaryZ. The datadesign is such that it allows for spellings that are correct according to an authority.
The other big problem is acceptance. Yes, it is the official spelling, but what if people and certainly big publishers do not accept it ? There is a new movement called "Witte spelling" that intends to create an alternative that is less confusing. This will result in a list of words spelled correctly according to this list. It results in a "green" and a "white" spelling. When we get the witte spelling as a resource, we can create a spell checker for Open Office, we can inform about the correct spelling according to the while spelling..
From a NPOV point of view, doing it in this way is problematic. The official spelling gets underrepresented, but how can we do it justice as it is proprietary? In several way the official spelling becomes less relevant..
If anything this is a great example that making what is supposed to be a standard proprietary, is a self defeating strategy.
Thanks,
GerardM
Thursday, February 16, 2006
Building a community of developers
Many needs for improvement arise from within the Wikimedia projects. These needs are typically taken care of by people who scratch their own itch. I am particularly interested in stuff when it helps me with the WiktionaryZ project. Other people have a need for Wikinews or WikiSpecies functionality.
One problem with the current model is that the developers are nominally all volunteers. This is when you analyze it no longer working it is also not true; the best developers are being snapped up by organizations and consequently are working either for interested parties or they are no longer available for MediaWiki work.
This means that it becomes more and more difficult to get things programmed. WiktionaryZ is an ambitious project. It needs available programmers and it needs people who can program and know other languages than just any "European" language and English. This need is felt more and more acutely.
I hope to develop contacts that I have made in Africa. A programmer that came recommended to me by someone who manages a great project in Swahili as best as he can. Erik did write this nice specification for something called InstantCommons. We hope them to develop this for us. When this works out, we have some new MediaWiki developers.
As we typically do, we discussed what to do when this is a success and, when we have a need for MORE developers.. Because of my interest in Iran (Farsi and Luri), I said that trying a similar would be a good idea. This got me in an interesting argument; Iran is with its present policies and president seen more and more as an enemy. Some people consider this to be a reason not to use the contacts that exists. From my perspective, this is not rocket science and it is about words and how they are used and understood. If anything we should collaborate on this.
Another strategy that we could adopt is having a competent developer, someone who is also a good communicator help students when they do termprojects related to MediaWiki. Any project related to MediaWiki. I think up to 50 projects could be handled given that a project takes half a year.. This would mean on average 100 students.. The benefit would exist in two ways; a lot of work can be done in this way and we probably would have a retention rate of something like 5% of the people as developers for a period of minimally a year.
Building a community of developers is essential. It is however not that easy :)
Thanks,
GerardM
Friday, February 10, 2006
What is a language
From a technical point of view, such a definition is beneficial. For all the major languages it is simple; it is obvious that there is a language code that describes them. For some languages there is a language code because people in the west take an interest, tlh is one code for one such language. For other languages it is more problematic, some of them have their code but that can make the problem worse. Some tools rely on the codes to be their and applicable; OmegaT, a CAT tool, for instance relies on the codes that exist in its programming environment. This is a serious problem because this programming environment supports ISO-639-1 and only with ISO-639-2 the code for the Neapolitan language became available. Consequently a translation tool does not support MANY languages. Even ISO-639-2 does not really help; the Kurdish language is acknowledged not to be a language; it is considered a language family that consists at least of three languages. These languages are acknowledged in the ISO/DIS-639-3.
While the ISO/DIS-639-3 is a huge improvement, it gets opposition from a few quarters. Many people, particularly developers of software, are of the opinion that some 8.000 languages is too much. Other people are of the opinion that the number of languages is not big enough. But also some languages that had support can be considered a dialect of another language like what is happening for the Twi language. How this will be appreciated by the people who speak Twi is anybodies guess. Twi is considered to be part of the Akan language, this article on the Akan language is indeed another example of the systematic lack of attention Africa gets.
For a CAT tool, it is really relevant that it allows its users to use the tool to its fullest potential. This does not mean that standards should not be supported, it means that multiple standards should be supported AND that you can introduce user defined languages as well.
Thanks,
GerardM
Wednesday, February 08, 2006
Gmail now with instance message functionality
When you receive as much mail as I do, most of it overwhelmingly from the same "source", mailinglists. Many people are subscribed to the same mailinglists and the reason why most people use Gmail... It would be cool if it were possible to identify mail addresses like the WMF mailing lists the point is that they can be stored differently. The content is not personal and it would be nice if it could be treated differently, it would be nice if mailing lists could be identified and stored separately.
The great news today is the new chat functionality that was added today. People who do not use skype or IRC but who have Gmail now can be chatted with. Two people I have communicated with for quite some time, now are available for a chat.. really powerful and guess what, one of them uses Google talk.. That was really sweet.
Thanks,
GerardM
Wednesday, February 01, 2006
MediaWiki is secure software
Many of the MediaWiki implementations allow everybody to create and change articles. This is a conscious decision, it is part of the formula and consequently this is not a problem from a security point of view. As a consequence the problem of maintaining quality content and preventing people vandalising the content, is a management problem. The tools to manage this problem are diverse but many tools that are considered security tools are usable.
Often vandals do not know that what they do is useless. Often people add links to all kinds of websites in order to increase their Google-rating. The MediaWiki software indicates to the Google crawler NOT to include external websites for its ratings.
Blocking IP-ranges and users because of persistent vandalism is one. Trusting logged in users more than anonymous users is another. There are many Wikimedia projects and all of them still have at present their own users. In Februari, it is planned to develop single signon for the Wikimedia projects. Single login has been on the wishlist of many of the people who are active on multiple projects.
With single login, in essence a management issue with security implications, it becomes feasible to use this as a stepping stone for the implimentation of security features that help with the management of vandalism.
The feature that I would love best is to differentiate the strength of authentication based on where a user comes from. When a user comes from a school with a history of vandalism, it makes sense not to allow anonymous edits. There are many of these types of soft security measures possible.
On mailing lists about Wikimedia, there was talk about a patch that allows for logging in users who authenticate themselves with OpenID. The interesting thing was that people had two issues with this; first it would not allow me to use my MediaWiki ID as an OpenID. The second is that to some extend OpenID is going to fit into the YADIS framework (Yet Another Decentralized Identity Interoperability System).
Yadis is interesting because it is linked to the eXtensible Resource Identifier or XRI, a standard that is developped by OASIS. It is also linked to the W3C (YASB - yet another standard body :) ).
In the end it comes back to standards; when the WMF would support twoway YADIS authentication, it makes for a VERY relevant implementation of security related functionality. This could provide for better management in the fight against vandalism. It is however important that it is a standard that we provide. That is why I am of the opinion that the WMF should support standards.
Thanks,
GerardM
Tuesday, January 31, 2006
Rule number one: You are wrong !
Today a nice standard called OLIF was pointed out to me. OLIF stands for Open Lexicon Interchange Format, it is an open standard for lexical/terminological data encoding. It is essentially a Free and Open standard, it has many illustrious and industrious backers. But it is a tad Eurocentric. It now has East Asian language support, but it is best at English, German, French, Spanish, Danish and Portuguese.
I would not want to dismiss a standard like OLIF when people are actively involved in a standard .. To be useful would be necessary to define how a standard is lacking, it might be that it is just a matter of some refining. It might be that the standard does consider things that I have not considered yet (and I do know that I do not consider everything all the time).
On Meta we are voting for a new Wikimedia project, this project is about standards .. Discussing standards but more importantly it is a conduit for working on those standards that do not fit the requirements that we have for standards in the Wikimedia Foundation's projects. Now there are two options: I am wrong or the standards will fit our needs like a glove.
Thanks,
GerardM
Saturday, January 28, 2006
About new words
If you know Dutch, you will love it. If there is more of this, also on other languages, on the Internet .. I would appreciate to learn about it :)
Thanks,
GerardM
Monday, January 23, 2006
Standing on the shoulders of giants
I have learned a few things. I am right when I understand that the location of words in a sentence is relevant. The TST-Centrale uses a fixed notation for it; I wonder how universal it is.. On the CD it is explained that people who use this content, can select the information they want to use and that this is part of the secret of its success. We hope that having the WiktionaryZ available as a database will serve in a similar way.
Reading back the first paragraph, I feel like a Tom Thumb. Every second noun is something to look up and it was hard work writing it. The great thing is, that when you have it in front of you, when you see it demonstrated it does make sense. When we collaborate on content and stucture, when we make WiktionaryZ something that is usefull because it has an application, we will have giants that help us little people make progress when we are allowed to stand on their shoulders and move with us forwards.
Thanks,
GerardM
Saturday, January 21, 2006
A discussion on trusted computing
From my perspective the biggest problem with trusted computing that I have is that it may be an open standard, but it is not a free standard. The specifications of the standard that the trusted computing group is available for organisations, it costs at least $1000,- and thereby excludes all these people that create wonderfull programs that people share. Because of this lack of openness it is has a fundamental problem. It makes me trust organisations that I not necessarily trust; why should I trust Yahoo, MSN and AOL as they demonstrate what I perceive as a lack of protection for the privacy of their clients? This trusted computing architecture does not allow me to trust my own software of the software created by a friend. It does not because I do not see how I can create software that will be trusted by my system.
Personally I do not think trusted computing is the equivalent of digital rights management. I am of the opinion that DRM leads to giving away rights that are mine.
Trusted computing does one other thing. I expect that it will take away much of the anonimity that is still with us on the Internet. This aspect is probably something that few people considered. My first clue came after I realised that it is a perfect tool to do vandal fighting on Wikipedia. It gives us a tool to more precisely know where these people come from and it can give us much better protection against this scurge. The other side of the coin is that when it provides us with more control, it will also give more control to those that I do not trust to use it wisely.
Thanks,
GerardM
Sweet nuggets of relevancy
It is only lacking in this one resource.. money.. If it does not find money it will pack it in. They are now finishing off software so that the project will be in the best shape it can be. I truly hope that the Kamusi project will find its way into a bright future. If you can help, please do :)
Thanks,
GerardM
Friday, January 20, 2006
etymonline.com
As it says on the information on the mainpage, they only have some 20.000 visitors, the size of the database is "only" 51,08 MB. My reaction to these problems are predictable; I would like to host this information in WiktionaryZ when we are ready for etymological content. There are however a few issues. These issues have to do with the perception of Wikipedia.. Let me quote:
-
"Approach Wikipedia about a partnership, or actually merge the site into Wikipedia. This is a painful option. In a sense, this site is the anti-Wikipedia. It is deliberately not open source. You'd understand why if you saw the regular stream of e-mails I get from people insisting that their own crackpot folk-etymology idea is absolutely correct. Such things can be based on some fervent politico-racial agenda, simple insanity, or "my French teacher in 10th grade told me."
A major reason this site exists is to serve as a template against which to measure people's best guesses and wacko theories. The whole Internet is a big Wikipedia; this site is a compilation of the most rigorous academic information."
In WiktionaryZ we have a need to address quality. In the current Wiktionary we have our crackpots. There is this loon who insists of there being a word called "exicornt". He even threatened admins in order to have this word exist in the Wiktionary.. :( We have to deal with these persons.
We could do something for etymonline. We can approach them. We can offer to host its content by importing their content (obviously with proper attribution), but we sure have to address the issues. Quality is important and we have to protect quality content for our own sake. When we do, and when are successful opportunities like this will be less problematic.
Thanks,
GerardM
Thursday, January 19, 2006
NPOV and sources (continued)
This is marvelous resource; there are some caveats. The on-line data is only from 1979 onwards while the resource started in 1932. The printed edition is available in the UNESCO library in Paris...
There are many possibilities with a resource like this.. It is indeed one of the resources that would benefit on a massive scale from a GOOGLE Book action. Obviously I would not care who does the scanning and who spends effort on digitizing this conten. It is however an extremely important resource and everyone would benefit if the data from 1932 to 1979 becomes publicly available as is the intention of this project.
There is more to learn about this project. When considering using it for presenting our sources in Wikimedia projects, we would like an API to refer to it.. I have not looked into it.
Thanks,
GerardM
Wednesday, January 18, 2006
What do you do when your computer is broken
What do you do when the WIFI router does not work anymore .. You use the ethernet option this router provides...
What do you do when the keyboard mapping is broken .. you do not use question marks ..
Life is a bitch.
Thanks,
GerardM
Monday, January 16, 2006
What to write about ..
Advertising
Yesterday Amgine wrote an “Advertising proposal. The idea is that we can advertise the services that the Wikimedia Foundation provides. Essential for good advertising and marketing is that you target your audience. With such an attitude comes the focus that would improve our content. The advertising would be done using an advertising server; the many people that have their own wiki, their own blog, could include advertisements from this server.
Wikidata
The Functionality that was originally conceived for Wikidata will end up in Mediawiki it self; this is both a blessing and a curse. The great thing is that it does signal that much of the functionality that we conceived is indeed relevant and it will bring functionality to all the wikis that use Mediawiki. The drawback is that it has implications for the design. It does complicate things at this stage, but on the other hand when we have a great instead of a good design from the start, the extra time and effort needed now may pay itself back in the long run.
Thanks,
GerardM
Sunday, January 15, 2006
Our blog
All these things are not new. It happened all the time.. Now we will make this process more visible; it is done to read this other blog, the WiktionaryZ.blogger.com blog of the commission.
The thing I am not sure about is to what extend I will continue blogging here and to what extend I will blog on the new blog.
Thanks,
GerardM
Saturday, January 14, 2006
Presenting: the Commission
There is one issue; people think that the problem with this project is that it is run by one person; me. It was once put to me was that that I represented a "truck factor" for the project. From my perspective, WiktionaryZ is certainly my dream, but it has never been my dream alone. I have worked hard to make WiktionaryZ happen, but I am not the only one that worked hard to make it happen this far. I came up with many ideas that made WiktionaryZ what it is, but not all ideas have been mine, all the ideas became WiktionaryZ because of the many conversations, e-mails and IRC chats about them.
From my perspective, there is this issue that I do not scale. There are things that amount to policy and it needs to be expressed because policy dictates technological choices and we are building the technology for WiktionaryZ. This implies that some policies have been set but it also implies that more questions will raise their head that do require an answer and require an answer quickly because we are building the software, the network, the connections now.
I did discuss this issue with Jimmy Wales, and he came up with the great idea to have a commission for projects like Wiktionary. This commission could have several funtions; it can act in a similar way as the chapters do for countries for projects. One role of the commission would be that the policies of a project would be consistent with the aims, the policies of the Wikimedia Foundation. Another equally important part would be to represent the community that will make WiktionaryZ its own. To do all this, members of such a commission have to be part of the discussions about the developing Wikidata and WiktionaryZ.
I like the idea and, I have asked several people to become part of an initial commission. They are all Wiktionarians (with one exception), they represent many Wiktionary projects and they are and have to be communicative; they do use Skype/VOIP and often can be found on IRC.
The commission:
Including Sabine Cretella is obvious. Sabine developed WiktionaryZ with me from the start. WiktionaryZ is a dream she has fostered for a long long time. Sabine is active on the Italian Wiktionary and is one of the initiators of the Neapolitan Wikipedia. Sabine is a professional translator and is known in this world as an evangelist of Open Source and Open Content.
Of all the Wiktionaries, English is the most relevant. Dvortygirl has been active there for a long time. Like me, she is an admin and she is well liked and respected for her work.
Gangleri is active on many wiktionaries. The thing I really appreciate is his involvement with right to left languages like Yiddish. On the Internet, these languages have their own issues, Gangleri is active in the Mozilla organisations to address several of these.
Yann, is active on the French, Hindi and Gujarati Wiktionary. He is also the treasurer of the French chapter. One of Yann's challenges is to help us get more people interested in the languages from India.
Erik is the one who is not into Wiktionary. The reason why he is invited is because he is the architect and realiser of Wikidata. This is the enabling technology for WiktionaryZ. Erik is also the realiser for WiktionaryZ. Wikidata has in WiktionaryZ its first application. As this is a truly big and complex project, many of the things that would hit Wikidata eventually, need to be addressed from the start. Erik is also important as a linking pin to the Mediawiki developers.
GerardM, if I need an introduction I invite you to read this blog.
Some policies or, a glossary of our policies:
Availability: "our data is to be made available through open standards and in a non-discriminatory manner"
Data design: "WiktionaryZ is implemented as a relational database. When information that is relevant in the context of WiktionaryZ cannot be added, we will try to ammend the design."
Full functionality: "WiktionaryZ needs to be able to include the information that is available in the Wiktionaries"
Partner: "a partner is an organisation that collaborates with us in the realisation of what we intend with WiktionaryZ"
Sponsor: "a person or organisation that donates money or content to the project or to the Foundation"
Success: "success is when people find an application for the WiktionaryZ data that we did not think off."
User Interface: "the interface should be in any language. We want this both for the Mediawiki and the WiktionaryZ user interface"
Thursday, January 12, 2006
Dialects within a language
JAVA uses ISO-639 for its language codes. The codes used is the ISO-639-1. Consequently the Neopolitan language is not known. OmegaT is an open source CAT tool, it uses the languages known to JAVA as the languages that it can translate.. So in order to translate to Neapolitan you have to pretend that it is a different language.. Not nice.. So the nice people of SUN were asked this and we have great expectations.
Today there was a request on Meta, the website about the Wikimedia Foundation's project, for a new Wikipedia. The request is for tarantino, it is considered a dialect, a dialect of Neapolitan. This request is problematic because there is not even an ISO-639 code. Consequently there is little chance of there being a wikipedia for created. Now, with the new namespace manager, it is possible to create a seperate namespace within the nap.wikipedia.org for the tarantino dialect. This is also a solution for the problematic request for a Lower Saxon wikipedia that will be in an orthography that is not German..
It is sobering to see that standards can enable and prevent things to happen. Good standards are vital and ISO/DIS 639-3 is a big move forward.
Thanks,
GerardM
Wednesday, January 11, 2006
Machine translation
To create a Machine Translation engine for these language, you would need all kinds of rules about the languages that you want to translate to. Now I would like to know these rules. Not so much to build Machine Translation engines but because they have this other application, one that is much closer to my heart, it is needed for software that teaches people languages. The Universität Bamberg is working on exactly this. They want to use the data of WiktionaryZ for this purpose. So if you have a nice sets of rules for them I would be obliged.
Thanks,
GerardM
The importance of statistics
As astonishing as the growth of Wikipedia is the apparent popularity of Wiktionary. According to Alexa Wiktionary is more popular than dictionary projects that have much more content. This is probably the effect that the association with Wikipedia brings us.
When you look at the traffic details for Wiktionary, the thing that strikes me most is the popularity of the Russian wiktionary. Such details points to the apparant strength of this project or to the need for Russian content. For me, this is relevant because it could be one reason why we concentrate resources on a given language.
As WiktionaryZ is a true Wiki project, there is no need to concentrate to much about the user interface. People WILL find what works and what does not work and to a large extend the user interface will evolve. However, at this moment we are thinking about the infrastructure of the project.
In the current infrastructure a resource is indicated by preceding the project name with the ISO-639 code; the Dutch wiktionary is therefore http://nl.wiktionary.org. For WiktionaryZ we do not have a separate database for each project. When we maintain this link into the mainpage for a language, we benefit in two ways; there is a main page for the whole project, there is a localised entry level per language and the statistics of Alexa remain relevant for some basic analysis of the demand of our project.
T
Tuesday, January 10, 2006
Transliterations
The current implementation of our database is the first itteration of what is going to be WiktionaryZ. Its intention started off as at least including all the information that is included in the Wiktionaries. So far I was against the inclusion of transliterations as we would want the translations anyway and, often a tranlation is in effect a transliteration. The case was made that this is too simple.
Several point were made:
- people find transliterations usefull
- many phrasebooks include them (their transliterations are often quite bad)
- having a "standard" transliteration helps because for names like معمر القذافي in excess of 20 different transliterations exist.
- transliteration should exist in addition to translations AND they are specific to a language
- there are standards for transliteration (eg for transliterating Russian into English)
Given my stance on what should be included in the database, I would say I want to have this. However, this is complicated by the fact that transliterations are language specific AND they are on the same level as pronunciations are.
Yes, this gentlemen knows how to find Njoe Jork, he studied some Afrikaans at one point in his life.. :)
Thanks,
GerardM
Just another day
- We have discussed that we are in danger of getting too much content. This may mean that we have to slowly absorb content and not make a big mess.
- There was someone new to me who came to Sabine and wanted to help us connect with someone we are already working with. The great thing is that this is a great person that can help the Wiki for Standards project along. He is also someone who could become relevant for WiktionaryZ
- Today I found this proposal called Wikimaps This is a proposal that was missing localisation, so I added it as a "must have" feature. When this is going to happen, many translators will be happy for a resource that helps with geographic data.
- I have been thinking more about partners. There is so much work and there are so many problems that will come our way and there are so many organisations that can help us manage. It would be folly not to collaborate and it would be ungrateful not to acknowledge these organisations.
- Quistnix told me that conversions are happening because messages about "namespaces". He is very active with interwiki links so he would notice these things.
- I received a mail about Maltese verbal morphology. This came with a request on how we could collaborate.
Thanks,
GerardM
Monday, January 09, 2006
What words do we want in WiktionaryZ
- The way copyright infringements are signalled in relation to lexicological content
- Do we want to emulate what others have done, or are we masters of our own destiny
- Where are the French, the Swahili, the Kannada and the words of all the other languages that are equally deserving attention
With WiktionaryZ we have the opportunity to use Wikipedia as our corpus. In Wikipedia we find the words as they are used today. When we concentrate our effort on these words, we provide added value to the information contained in Wikipedia and at the same time Wikipedia adds value to WiktionaryZ because it allows us to show words in context. In Wikipedia we have people from many countries that contribute, they do use words that are normal in their locale. With the "OED and AHD" we do not necessarily get these words.
It is not a bad thing that people interested in English concentrate on English. As such I welcome this effort. However, in my opinion the emphasis is too much on main languages like English. In these languages it is hard for WiktionaryZ to become relevant. To become relevant we have to do things that others do not. Relevancy can be gained in many ways; translation to minor languages is a way for some, counting the characters in a word makes is a way for others.
From my perspective, we become relevant by harbouring communities, special interest groups and allow them to make WiktionaryZ their project. When we maintain our core values of Freedom, of inclusiveness, of non-discriminatory access, of open standards we will be relevant to some if not to all.
Thanks,
GerardM
Sunday, January 08, 2006
Welcome Sabine :)
When we are going to have single login, people will know that Sabine is not a newby. So what will happen when you are new to a project, would that mean that when you are new to the project you will not be welcommed ?? Sabine and I found it funny that she was welcomed then again it is the charm of our project that you are welcomed.. What is funny is this..
Thanks,
GerardM
Mediawiki is very much a collaborative effort
The biggest improvement for the production of statistics has been the policy that the dumps of the Wikimedia Foundation are in a XML format. This provides a much more stable basis for the production of statistics. With the advent of Wikidata and single login, I worry about how this will affect these popular statistics.
It was great to have Erik reassure me that we will cross that bridge when we meet it, "because we have always done that". Issues that may arise are: currently we do not have Wikimedia wide statistics, with single login implemented we can. Wikidata projects will be .. where and will probably not count for the project that the data has been entered for. With the introduction of the namespace manager in Mediawiki 1.6, some of the assumption in the statistics may have to be revisited..
Thanks,
GerardM
NPOV and sources
Let me be clear about my position; I am all in favour of providing sources with articles. However, for every controversial subject you can find literature that "proves" all the crackpot ideas that are floating around. To complicate things even more, literature available in one country or language is not all the literature that exists on a subject. It is therefore my position that you do not prove anything by providing sources, you only prove that there are sources that helped us come up with a particular article and that it raises the standard of quality.
Yesterday, there was a New Year meeting of Wikipedians in Rotterdam, and as is usual at such meetings all kinds of everything were discussed. Including the problems with sources. During the discussion we came up with the following:
We need a central database to do away with the current interwiki links. When we have such a database, we should have all sources for the same subject, never mind what Wikipedia it is, in there as well. It helps people find sources but also when there is a difference in the view taken between the different Wikipedias, the sources can be compared and it will prove to be a usefull instrument to deal with cultural bias.
The thing that triggered this idea was that Oscar had bought a book about the Amazingh. On the Dutch Wikipedia there was a guy who quoted sources that noone of us could read as it was in the Berber language. I was really pleased that Oscar took the effort to buy this because it shows very much his good faith, our good faith. By having all sources for the same subject in one place, we would show similar good faith and, it would be a really powerfull tool to remove much cultural bias from the Wikipedias.
Thanks,
GerardM
Friday, January 06, 2006
Portal pages in WiktionaryZ
When people want to do this kind of research, it makes sense to have a place for it in WiktionaryZ as well. It dawned on me that I have not really considered many of the User Interface questions that will come up. Then again some things are obvious. People will not only want to select the language that they see. There will be a need for a portal page for each language. This could be the kernel for a portal page for the English language (from the Dutch Wiktionary).
With starter portals like this, it can be expanded in many ways. Many internal (to the Wikimedia Foundation) and external resources links can be added and give users a rich experience. The Main page of WiktionaryZ would therefore be similar to the http://wiktionary.org website particularly to point people in the right direction.
Thanks,
GerardM
Thursday, January 05, 2006
Angela's report
The fun will start when English, South African, Australian, New Zealand, Canadian and US-American legal terms end up together. The only way this can work is by having distinct terms clearly associated with a specific legal glossary for a legislation.
The same report mentions a biblical Hebrew dictionary, it seems that people working on such a resource want to have it in a collaborative environment.
Both are examples of the pent up demand that exists for a resource like WiktionaryZ.
Thanks,
GerardM
Wednesday, January 04, 2006
Organizing the Ultimate Wiktionary project
We ended up with a design that will allow for a lot of refinement. However many people who looked at it think it has great potential. Ultimate Wiktionary, the project, is our dream. As we worked on it for the last year and a half we understood more and more the pent up demand that exists for information that can be stored in our project.
Our dreams are big. We want to realise them. We are fortunate as we are associated with the best organisation to make this happen, the Wikimedia Foundation. It has a great reputation, it has a great community and what they do with Wikipedia is astounding.
From an organisational point of view, the Wikipedia project will be different from the WiktionaryZ project. Wikipedia is community driven; they create the data they finance the project. WiktionaryZ will be different because much of the data that will be included already exists. Many organisations struggle while maintaining their resources. For WiktionaryZ the opportunity exists to focus all this energy in one place.
The development of WiktionaryZ was made possible by organisations supporting our effort. Kennisnet, a WMF partner, provided the initial investment. The Universität Bamberg was the second organisation to help. More work needs to be done and there are more organisations that are willing to collaborate technically and who are willing to share their resources to make WiktionaryZ happen.
WiktionaryZ is going to be big. I made the bet that we will need in the first year of full-featured production two hundred thousand EURO (not US$) in servers. There are people that have told me that my “guestimation” is on the low end.. :)
People that work on content are and will be attributed in the normal way; it can and will be found in the history of the content. Organisations that prove to be partners of our project could be credited on the left hand side and may end up under the toolbox. A link will refer to a page about our partner.
Our project is as much about collaboration as any of the other Wikimedia projects. There is however no other project where organisations will play such an important role. This calls for a different way of organisation their effort. I propose therefore to combine these organisations in a consortium that will be the focal point for the contributions of organisations.
The WiktionaryZ consortium will have two functions; managing the collaboration of organisations and finding the resources to make WiktionaryZ possible.
Thanks,
GerardM