Sunday, August 20, 2006

Latin and species

We have an ever growing list of languages in WiktionaryZ. There are always good reasons why we add a specific language; it has a nice script or there is someone interested in adding content or we have a potential partner with an interest in THAT language.

It was suggested to me to have Latin. The motivation was; I have this list of birds and would it not make sense.. It does make sense on one level and, on another it does not. The Latin used to come up with all these taxonomical names is not necessarily the kind of Latin that helps you learn the language.

When we want to include THAT kind of Latin, it makes sense to include the taxonomical relations and attributes that makes taxonomy the science that it is. Considering this, it will need Wikiauthors for it's publication data. It will need some specific database functionality to make this feasible..

I do want Latin but I doubt if this is a good moment to include it.

Thanks,
GerardM

Friday, August 11, 2006

After Wikimania ... semantic mediawiki

I had a great time at Wikimania, and I am now slowly but surely trying to get things organised. This is not a trivial thing.. loads of people.. loads of interest in what we are doing with WiktionaryZ. The need for becoming organised is becoming increasingly important.. So how do I this, what tool to use.

When I was in Rome I met Denny. He demonstrated how well Semantic MediaWiki can be used for personal use. I was really impressed. Denny installed all the necessary bits on my computer. And it does do many of the things that I need. It allows me to create multiple lists associated with a subject. For instance the organisations involved and the persons involved with a project..

The more I see it, the more impressed I am. The only potential problem I see is that people have to learn more wiki-syntax. As it is so powerful, I would not mind to include semantic mediawiki in WiktionaryZ.

Wednesday, August 09, 2006

Yochai Benkler: The Wealth of Networks

At Wikimania there were many presentations. Some were great, some were awesome. It is really hard to assess how influential many of these will be. There are ways whereby presentations may become relevant.
  • The message was heard for a first time by a public.
  • The message was told for the first time.
  • The discussion following a presentation brought new insights.
  • The presentation raises questions.
I have truly enjoyed Wikimania. For me the presentation of Yochai Benkler was intriguing. It is about the way Internet is changing how information fits into society. What I understand from the presentation is, that business as usual has had its day and, we are working on what the new model will be for the future.

As I did not grasp what the new role for organizations will be in this brave new world, I bought the book. I bought the book because I strongly believe that there will be a role to play for organizations and if there is to be a recipe for including organizations, businesses I want to have it. If there is no such recipe ..

I have browsed the book so far, many things I do not grasp. In discussions I had earlier in the week I learned how different the notion of "liberal" is depending on the context. As this is also a strong theme in the book, I will be struggling. However, it is fun.

Thanks,
GerardM

Sunday, July 30, 2006

WiktionaryZ and attributes

At this time, WiktionaryZ is able to have text based attributes. The implementation allows us to have attributes on the level of the DefinedMeaning. In plain English it means that we can have a free form text without any wiki-syntax. It does not do much for us at this moment. It does not allow us to have sample sentences (they are on the SynTrans level), it doe not allow us to indicate part of speech, inflection or gender (they are on the SynTrans level).

The good thing is that it is an indicator of things that may one day make our day. The thing however is, that it does not necessarily have our highest priority. At this moment our priority is to get our versioning working. We need to have historic data, we need to have improved information in our recent changes.. This is what we need badly.

We need it badly because this is the core functionality of WiktionaryZ or Wikidata. It needs to be done. It needs to be done well. It needs to be done before we add even more complicated functionality in WiktionaryZ.

There is a lot of work that needs to be done. We want to have proper support for terminology, we need to include the notion of "domains" for that. We want to have proper support for lexicology, we need to include the notion of language dependent attributes for that. We need to have relationtypes that can only be chosen in the correct context. The context being either the language or the domain.

All these things will happen. They will happen in a way that is consistent with the amount of resources available. This is why we have a need for involvement. This involvement can come in many different ways from from many different places providing many different angles to improve our data. The key thing is that we will maintain our architectural integrity. It is completely unacceptable to build all kinds of wished for functionality while the technical foundation has not been laid.

The collaboration that became WiktionaryZ will start it's third year at the end of August. There were many reasons why it has taken this long. What it brought us is a design and a philosophy that may work. The third year will be the year where I expect that we will have our full functionality.. When we do, things will evolve more quickly.

Thanks,
GerardM

Tuesday, July 25, 2006

The operational definition meets the DefinedMeaning

In WiktionaryZ, we aim to have definitions for all the concepts that are to do with an expression in a language. These definitions have to be good, they have to express well what the concept is. We have great definitions, operational definitions, when more than 90% of the people correctly identify the concepts given the definition in a corpus.

When such an operational meaning cannot be correctly identified by 90%, it means that all the definitions of the different concepts are suspect. It may mean that too many concepts have been identified; certainly more work needs to be done to define things better.

In WiktionaryZ, many DefinedMeanings may exist where it has been indicated that the definition does not define the concept really well. These concepts are close but no cigar; they should not be taken into account when it is determined if the Definitions are indeed operational definitions.

The question is, how do you then identify the quality of the translations and the usability of the synonyms.

Thanks,
GerardM

Sunday, July 23, 2006

Being too busy

The last month has been a rollercoaster; the new functionality of WiktionaryZ that we now have is awesome. The whole idea of the [[DefinedMeaning]] starts to make sense. People are collaborating on the same data. The idea that information can be shared and that it only needs to be added once is now a reality.

The thing that amazes me is that I am so busy making sure that everything works well, that we have some crucial data. Things like the Swadesh lists and the the list of 1000 basic English words are really important because these are basic words, they demonstrate best the merits of the concept. And I am really pleased with how things are progressing.

The thing that annoys me is that there is so much work that I would like to get done.. Some of it is plainly not for me to do. Other stuff like writing on the blog is very much for me to do. I may have to be even more selective what I do. Making these choices is hard.. Then I am glad I am in this position.. So many great things are happening and I am part of it :)

Thanks,
GerardM

Tuesday, July 11, 2006

Water; at least three DefinedMeanings

When the word "water" is considered, people say it is this liquid that people drink and also that it is this chemical known as H2O. Well, actually these two things are not the same. The chemical, would be closest to distilled water and, it is not healthy to drink. The water you drink has traces of all kinds of chemicals like salts in it. This what makes for it to be good to drink. Another use for the word "water" is to indicate a body of water where you can swim, raft or boat on. In essence it is short for open water and as such it can be both salt and fresh water.


It has been said before by many, it is the simple words where "everyone" understands what is ment which are the hardest to define.

Thanks,
GerardM

Saturday, July 08, 2006

Just a word

Working on projects like WiktionaryZ is a job do take a lot of my time. Often I would like to add just a few words to my home project, the Dutch Wiktionary. Today, I felt a need to add a word, record the pronunciation and translations to the English word for the same phenomenon. The word is uitzaaiing.

In Dutch the word has a clear agricultural background. It refers to how a weed, once it is established seeds itself to neighboring areas. It is one of those word that you hate to hear in another domain. Today I did. There is so little that I can do, I feel sad and can only hope for the best for Scott.

Thanks,
Gerard

Sunday, July 02, 2006

A dubious record

I run for the Wiktionary projects the pywikipedia interwiki bot. This is a program that finds articles by the same name in different Wiktionary projects and creates a link to the projects that share this article. I run it for quite some time now and as the projects grow bigger, they share more articles.

The bot is quite nice, it can run autonomously and it updates some 30 wiktionaries at the same time. For the smaller wiktionaries, I run it specifically for a project every now and then.

Yesterday I chalked up the 200.000th edit for the English Wiktionary. It makes it nor unrealistic to think that I have some 600.000 edits on all the wiktionaries. There are six instances of the bot that run at any one time, they run on different projects. I think it is a dubious record because it does not bring me any happiness; it only indicates that there is information about a word that is written in the same way. I doubt very much that anybody does anything with it.

For your amusement; there are MANY English words on the Chinese Wiktionary that cannot yet be found in the English Wiktionary .. :)

Thanks,
GerardM

Saturday, July 01, 2006

What is your mother tongue

Sabine her kids speak primarily Italian, they are living in the area where Neapolitan is spoken but Sabine is German. Sabine talks frequently in both Italian and German to her kids and when she gets angry she turns to Neapolitan, it can be really expressive .. :)

Now what is the mother tongue of Sabine's kids ? They speak primarily Italian...

In WiktionaryZ we have people that indicate that their mother tongue is zho or Chinese. According to Ethnologue Chinese is a macrolanguage. This implies that Chinese cannot be a mother tongue, one of the 13 languages Chinese is divided in can only be the mother tongue. This is a potential hot potato when people equate Chinese with the country and not the language.

WiktionaryZ is about languages and only about languages.

It can also be understood differently, they may mean that Chinese is the first written language that they learned. However if I understand things well, when people talk about the Chinese written language, it is actually Mandarin. For Yue for instance, there is a need for additional characters that are in one of the later versions of the UNICODE. This is however not what I would consider a mother tongue. A mother tongue is the language that you learned from your mother. Writing is what you learn at school.

Thanks,
GerardM

Saturday, June 10, 2006

Punjabi and what IS that script

On the WiktionaryZ main page, we have a list of languages, they point to "portal" pages for those languages. It is quite clear that a project like WiktionaryZ has to take the different scripts into account that a language may manifest itself in. After a lot of head scratching, I created a link to "cmn-Hans" and "cmn-Hans", to indicate that there is Mandarin both in a simplified and a traditional script. One other reason, it looks more organisms this way.

I then tried my hand at the Punjabi language. Punjabi is written in two scripts, and I guessed wrong trying to identify them. One was indeed an Indic script, ਪੰਜਾਬੀ is written in the GurmukhÄ« script while پنجابی is written in the Shahmukhi script. Shahmukhi is indeed an Arab script but it is not the Arab script. In order to properly identify these words, I looked them up at Unicode where there is a nice list of the ISO-15924 script codes. Gurmuki has Guru as its code and Shahmukhi .. is absent.

It is probably pretty safe to indicate it as Arab, but when my information says that it is not, it is indeed problematic. I could also have it as an uncoded script. The problem with these standards is that they work up to a point. The point is what are they there to do.

When you write Dutch the standard Latin script is used, however that leaves out one character and consequently all word processors capitalise the ij wrong, it should be IJ and not Ij.. I think a similar thing is happening with the Shamukhi script. It is assumed to be Arabic but the style of the glyphs is different. I think it is just one of those things that may change in the future.

I think I will indicate it to be Arab for now .. :)

Thanks,
GerardM


Sunday, June 04, 2006

What to do next

We have in our GEMET data our first data online. We have already planned the import of the languages that are in ISO-639-3 and allow for the translation of these names to other languages. We will also extend the number of languages that can edit. Given that we are in pre-alpha and that we are at this moment still very much a development project, we include all languages that showed some activity towards localizing the MediaWiki user interface.

The question is what to do next. We cannot and should not rush the development but we do have a host of data that we could import. It could be other thesauri, glossaries or ontologies. It could be a long list of Expressions in a given language to populate a spell checker. We could import the data resulting from Duesentrieb's Wikiword application, this could give us a link to Wikipedia articles.

When we had more active collaborating developers, we could consider doing the import and export routine or we could start working on inflections.

What would you consider the next bit of data to import after the languages? What would you start programming on given where we are at this moment?

Thanks,
GerardM

Wednesday, May 31, 2006

What use is a community for others ..

I was at the LREC 2006 conference in Genoa, and one recurring theme was the use of software because there are not enough people to do things manually. Some things a computer can do well, some things a computer does not so well. You are often presented with a percentage where the computer is off from what a human would do.

One presentation, the price winners presentation at the end of the conference mentioned a nice scheme were two concepts were compared and the question was, to what extend is the first concept associated with a second. People doing this are trained with a first set of concepts, they are then asked to do a further set of concepts to see to what extend they have learned things and then .. They are off. This worked really well but, you need a large group of volunteers or you use a computer. The computer did either a good job or gave COMPLETELY different answers from what a human would do (these are the things to watch our for in a Turing test).

My idea is, when WiktionaryZ gets itself a large community of people interested in languages, it would also be natural to ask this community if they are interested in helping out with research. One strategy would be to have just one person check a machine derived result, when there is a discrepancy, have some more people look at it ..

Another interesting experiment would be to test the difference between the different groups of users of English including people who use English as a second language.. What do you think ???

Thanks,
GerardM

Saturday, May 20, 2006

Languages, dialects and orthographies

When a text is known to be in a certain language, and this language is more or less familiar to a person, this text may be meaningful. I have had some French classes and when I am in Italy there is quite a lot that I can understand. An automated process cannot do this; it helps quite a lot when a text has Meta-data that indicates what language, dialect or script it is.

One of the things that makes sense to be aware of, is what orthography a text, phrase or word is in. It is definitely something that is in a class of its own and it matters when text is to be understood in an automated way. Languages do change over time and, the recognized correct orthography changes to reflect this. The German and Dutch language both have had their fair chair of changes. The functional design of WiktionaryZ has always had a place to indicate that a given spelling is dated. The way we will export the WiktionaryZ data will be by using standards like TBX, LMF maybe RDS, SKOS or something different but standard. The problem is; how do we indicate that a given word needs to be spelled different since a given date ?

Thanks,
GerardM

Friday, May 05, 2006

Languagecodes on the Wikimedia Foundation

Hoi,
In the past we started using the ISO-639-1 codes for indicating languages. This list was extended with the ISO-639-2 list because it was too limited. Even so there were issues with the list and they were augmented with codes from Ethnologue. This was not often considered not enough so we created our own codes.

Now there are the ISO-639-3 codes. They have a provisional status because there will be even further extentions of the list but it highlights one thing. Us using our "own" codes is really problematic. I give you some examples and I also indicate to you why many arguments used in the discussion on new languages are wrong.

The ksh.wikipedia is for something called Ripuarian. The ksh code is for the Kölsch language. Ripuarian is considered to be a language family. The consequence is that there IS no single Ripuarian orthography, language or culture

The als.wikipedia is called Allemanish. According to the ISO-639 these are four languages. The problem is that the als code is used for the main Albanian language. The code for Albanian sq (ISO-639-3 sqi) is also considered a languagefamily. The two main variants are the Albanian and the Kosovar languages. Two other members of the Albanian language family are spoken in Italy and Macedonia

Some of the most heated discussions on the request for new projects are about the status of a language; is it a language a dialect and often the arguments are of a political nature. The inclusion of languages in the ISO-639 has been political in the past. With ISO-639-3 many of these arguments have an answer with the many new language codes that have been created.

The result is that we have wikipedias like the ku the fa, the sq, wikipedia where the language is now considered a language family and where a request can be made for recognition of a language that is part of that languagefamily. There are more projects like that that I have not identified yet.

Another "nice" situation is the Low Saxon nds wikipedia. When you look at what Ethnologue has to say about the Lower Saxon language family than you get the impression to what extend there is not really something called Low Saxon in the Netherlands. The varieties of Low Saxon of the Netherlands are all there.. They are named, have there codes..

The point that I am raising is, languages are a mess. The codes for our projects are as a consequence a mess. The procedures for new projects are a mess because of the politics and the codes we have come up with in the past. What we need are better quidelines what the relation is between the ISO-639 codes. If the WMF says it uses the ISO-639 codes the codes must be in use or they must be clearly different.

Last but not least, according to the terms of use, we are not allowed to extend the codes in the way that we do.

Thanks,
GerardM

Friday, April 28, 2006

and now for some other news ..

I read on the BBC news website that the UN is cutting in half the daily rations of the fugitives in the Darfur region due to severe funding shortfall .....

From May the ration will be half of what the minimum amount required for each day.

They are starving and they are the lucky ones ......

Thursday, April 27, 2006

Collaboration

The classic business model for lexicology, terminology and thesauri is .... create a dictionary and keep it to yourself. Protecting this investment is difficult; they are facts so they are protected as a collection of data. The classical way of "proving" that the collection is "stolen", is by adding some nonsense words or have some other bogus data as part of your collection.

The open / free business model for lexicology, terminology and thesauri is .... work together on a stellar resource and make the data available to everybody. With the data available to everybody, the big question is how to achieve the best result. For an open / free resource, the best way is by providing the data in a standard way. The first standard I want WiktionaryZ to use is the "TermBase eXchange" or TBX standard. Given what we do in WiktionaryZ, LMF and SKOZ are two other great standards.

Why use standards? Simple, our definition of success is: "when people find a use for our data we did not think of". By providing the data in a standard way, it will be available in a stable way and as a result it will be more easy for people to make use of our data. It will be easier for WiktionaryZ to become a success.

Yes, I love our "competition", but I will love them to bits when they want to be as relevant as we want to be.

Thanks,
GerardM

Monday, April 17, 2006

What to do with stuff that is good but not standard

What to do if you can get a lot of content that is good in some respects and lousy in others. German uses different characters then the English language and there are ways to indicate an Ä/Ä, Ö–/ö, Üœ/ü or ߟ. When you are using German these characters should be used and not an ue for instance. So when should we accept content with German that is non standard ?

I have been thinking about this for a few days. The answer for me became obvious; it is in the database. When we have the MisSpelling table, we can have the community identify the words that should have an Umlaut. With proper logic the representation for our public will be the proper German with the umlauts.. But the first thing is to have the MisSpellings..

Thanks,
GerardM

Friday, April 14, 2006

Diana posted an answer ..

Diana posted an answer to an entry of this blog. I know she did because I was send a message to inform me of the fact. I was quite happy with her message, she wants to get into contact with me but I do not know how to get into contact with her.

I am GerardM .. :) You can find me on the WiktionaryZ site. I am very happy to talk about collaboration. I am happy to remind everyone how we define success for the project: Success is when people find an application of our data that we did not consider in the first place..

Thanks,
GerardM

Tuesday, April 04, 2006

Some of the best things in life are free

Last week I went to the Berlin 4 Open Access - From Promise to Practice conference in Golm Germany. For me it was an education. The really big thing that I now appreciate even more than before is the extend science is prevented from being science because of restrictive practices.

Typically something can be called scientific when the conclusions are arrived at in a methodical way and, the method is repeatable. This is exactly what the Open Access movement wants to bring back. In order to do this they have to wrest away the restrictions that copyright has put on scientific data far too long. A lot of bad science is the result of these restrictions a lot of wastage is the result of these restrictions.

What I learned is that many superb resources are becoming available to the world as a consequence of this movement. Open Access is a rich tapestry with many threats in many fabrics of many colours.

If there was one thing disappointing it was the lack of awareness of licenses. It is a CC license.. was the answer and people applauded. Well, Creative Commons has great licenses but it is a bit like Animal Farm; all CC licenses are Free but some are more Free than others ...

Thanks,
GerardM