Showing posts with label CLDR. Show all posts
Showing posts with label CLDR. Show all posts

Saturday, February 02, 2013

#CLDR gets the sorting right


When you sort, order will be created in the predetermined way. Another word for such a predetermined way is called the "collation order". When you sort tea, you make sure that only the tea leaves of sufficient quality are left. The characters in the words determine where the word can be found in a sorted list.

The collation order is a standard and, the CLDR is the name of the standard. Unicode, the organisation behind the CLDR has made a big change in the order. From now on, the character of the script of the language take precedence over the Latin script.

This change affects many languages and, there is a document mentioning them all.
Thanks,
      GerardM

Thursday, May 31, 2012

#WMDEVDAYS - #Wikidata

Presentations on Wikidata have started. There are enough thought provoking ideas here. Daniel Kintzler presents on the interlanguage links and they did think on how to make information stored in the info-boxes.

Consider a Wikipedia with less then 1000 articles. It is extremely likely that they do not have information about pope John III for instance. When you analyse this English info-box, it has 9 labels. Once these are translated, much of the existing data can just be presented. Date formats can be altered as needed based on the CLDR data. The names may need to be changed based on how a person is known; Benedictus is known as Benedict in English.

Obviously once all this information is validated, you can either create an article or you can provide the template as part of the "not found" information. An other option is to create the article when all the information is validated and show an incomplete info-box when the box needs more work.

We want to share the sum of all knowledge and making use of the info-boxes has potential.
Thanks,
    GerardM

Tuesday, May 15, 2012

#CLDR will know language names in #Esperanto

#MediaWiki uses the language names as defined in the CLDR. It is therefore important that people compile a list of the translations for their language make them available to be used as the standard translation.

Arno did exactly that. The list he created contains all the codes as used for Wikipedia combined with the translation in Esperanto. It is an important effort and it would be great when we have such a list for all the other Wikipedia languages as well.

As standards are standards, we were asked to provide the list in an XML format. It took some pastes and find and replaces and it looks good. The only problem is that some of the codes used are not standard codes. Several codes have been removed, "als" for instance is the code for Albanian Tosk not Alleman. There may be some other "language codes" in there that are not recognised in a standard.

When your language can do with additional translations, please follow the Esperanto example and provide a list of translations in your language.
Thanks,
      GerardM

Thursday, May 03, 2012

A #CLDR walkthrough

For #MediaWiki, the CLDR information is important. Sadly for many of the languages supported in a Wikimedia Foundation project the information is not (yet) available. Several things are needed;
  • People who know their language well enough to enter the data
  • People who know their language well enough to verify the data entered
These people exist for any language. The question is what does it take for people to enter the data. For the Asturian language, a language from Spain, the data is now being entered. 

One of the things that may help is instructional material and, it is quite wonderful that an instruction video has just been released. To quote the message announcing the material:
A new 52-minute walkthrough video is now available, showing how to use the CLDR Survey Tool to enter data, and prioritize your work. The video and explanatory material are available on this CLDR site page.

So please watch the video and do what you can do for your language.
Thanks,
     GerardM

Wednesday, April 25, 2012

Supporting #Asturian in the #CLDR II

It is that time when people CAN support their language and enter data about their language in the CLDR. Last year, collecting the necessary data for Asturian did not happen within the set amount of time.

This year another try will be made to find and enter the data for Asturian. It is good news. I hope to learn about more languages and locales that will be entered.
Thanks,
     GerardM

Sunday, April 15, 2012

Supporting #Asturian in the #CLDR

Finding the data to support a language in the CLDR can be a struggle. The core requirements are only a few so how hard can it be...
  1. (04) Exemplar sets: main, auxiliary, index, punctuation. [main/xxx.xml]
  2. (02) Orientation (bidi writing systems only) [main/xxx.xml]
  3. (01) Plural rules [supplemental/plurals.xml]
  4. (01) Default content script and region (normally: normally country with largest population using that language, and normal script for that).  [supplemental/supplementalMetadata.xml]
  5. (N) Verify the country data ( i.e. which territories in which the language is spoken enough to create a locale ) [supplemental/supplementalData.xml]
  6. *(N) Romanization table (non-Latin writing systems only) [spreadsheet, we'll translate into transforms/xxx-en.xml]
When you read this, the text indicating what the initial requirements are, it becomes quite obvious why this process has such a bad reputation. 
  • It is not clear nor relevant where the data provided ends up in an xml format
  • Orientation is very much an aspect of a script, not of a language nor of a locale.
  • In the survey tool it is Esperanto that proves that a language may not fit into a locale anyway.
  • Romanisation is stated as a requirement. It is however not obvious at all that every script or language has ever been romanised in a standardised way and why this might keep a language out of the standard
In the past there has been an attempt to provide information for the Asturian language to the CLDR. The good news is that there is documentation on why it failed. The problem was that when you establish data about a language, you need to be certain. Four Unicode characters ('Ḥ ḥ Ḷ ḷ') are used for writing the Asturian language properly and the literature on the subject was not consistent. This issue was resolved but it took more time then was available in the CLDR time box.

The Asturian example proves that getting data ready for a standard takes time. The practice of closing a request because the data was not provided within a set amount of time is what stopped people dead in their tracks. We can only hope that Asturians will find what it takes to get support for their language in this time box.
Thanks,
      GerardM

Thursday, April 12, 2012

#CLDR is not used solely for locales

#Unicode is best known for its architecture for the digital representation of the characters off scripts. The Unicode consortium also hosts the "common locale data repository" project also known as the CLDR. In this repository you can find how languages are used in different areas. You will find for instance what currency symbol is used. There is also information on what direction is the language written in, the way numbers and dates are written.

Many applications rely on this data. Several Open Source word processors use this data to enable a language for editing. While it is great that this data is used it is problematic because not all languages are spoken let alone in locales.

One great example is Ancient Greek; this language is not a living language but it is taught in schools all over the world. Students are doing their homework and what they write, it is certainly not modern Greek. From a technical perspective, it is only correct when the meta-data of such documents indicates that it is Ancient Greek.

When Ancient Greek and other extinct languages are supported in word processors, surviving texts can be written with modern tools. These transcribed texts will by default have correct meta data and it will be easier to find them when they are placed on the Internet.

For this to happen, the CLDR either embraces that it is used to enable languages in word processors or the word processors who currently use the CLDR allow for alternate sources of primary data.
Thanks.
     GerardM

Sunday, March 11, 2012

The web is Multilingual - #multilingweb

Presenting at the conference about the Multilingual Web will be fun. When you read what the conference is about, what they ask presenters to include in their presentations is interesting:
  • existing best practices and/or standards that are relevant
  • new standards and best practices that are currently in development
  • gaps that are not covered by best practices and/or standards
In so many ways, what we do is implement the best practices as we know them. We are establishing best practices and are running into the gaps of the standards regularly because no other project supports the 412 languages that have an existing Wikipedia or are requesting a Wikipedia

It is great for us to have two people at this conference; we will learn a lot from the other presenters and from the people who attend. We expect that many best practices are set into a professional environment. Our environment consists of dedicated volunteers. The monthly update to our community of localisers at translatewik.net has 4600 recipients. Our puzzle will be to adapt what we learn for our setting.

We want to learn about translation work flow, we want to discuss what to do about languages that are not yet supported in the CLDR. Most of all we want to learn what we do not know, our blind spots. 
Thanks,
     GerardM


Wednesday, February 22, 2012

#Unicode #CLDR How to fix a wrong name

One of the languages in the Incubator is Mapudungun. It is spoken in Chile and Argentina and many of the Mapuche people are offended by the English name used for their language. The result is that they refuse to have anything to do with the effort to have a Wikipedia for their language.

The Wikipedia article on the Mapuche language mentions the name that should not be used and it states: "The latter was the name given to the Mapuche by the Spaniards but nowadays both the Mapuche and others avoid this usage". The question is; why does it say Araucanian, where does it come from and how can it be fixed.


MediaWiki uses the CLDR for such data and with some digging the record that has Araucanian can be found. The ISO-639 standard does not have it as its name so it can be changed. The next question is, who is going to do this. There are several issues:

  • the CLDR has only windows of opportunity when the data can be changed
  • actually changing things at the CLDR is an "experience"
  • we do not have anyone volunteering to join the language support team for English
  • we do need this change to be in the CLDR for this issue to go away
Thanks,
      GerardM

Wednesday, February 08, 2012

Plural rules OK

#MediaWiki aims to support over 300 languages. Not all languages are equal in the way they express things like a plural. The rule for English is easy: one, many. One apple many apples never mind how many, it will be apples.

Other languages express plurals different. How they are expressed is defined in the CLDR standard. They use a formula and consequently languages that express plural in the same way share the same formula.
<pluralRules locales="hr ru sr uk">
<pluralRules count="one">n mod 10 is 1 and n mod 100 is not 11</pluralRule>
<pluralRules count="few">n mod 10 in 2..4 and n mod 100 not in 12..14</pluralRule>
</pluralRules>
These formulas, when known for a language, will be used in MediaWiki when generating messages including numbers. As we already know the language, we just have to look up how a number fits what rule. Using these formulas helps because it prevents us from having to write separate code for each and every language and, when the rules become known for more languages, they will just fit in.

The challenge is to ensure that the CLDR will support every one of the existing 6000+ languages of which we at present only support a twentieth. 
Thanks,
      GerardM

Monday, January 30, 2012

#Wikimedia celebrating Februari 21

Every year the UN celebrates International Mother Language day. It is a moment to reflect on the many languages spoken and what we can do for their future in our society.

Some say that having a Wikipedia is one of the best ways of demonstrating the vitality of a language.  In a way it is. It is a place where a language is used to articulate all the knowledge in that language. When a subject is explained that is rather foreign to that language it is a challenge to cope. People manage by inventing new words or by borrowing from other languages.

Technically, we do need to know details about the grammar of the language, we need to know what script is used. These details are the same details that enable software developers to support a language in their application. This includes word processors, browsers. Everything you need to be active in this modern and increasingly digital world.

The Wikimedia Localisation team will host another "Office hours" on Freenode on February 21 2012 at 18.00 UTC. It will be a perfect time to discuss everything we do to support languages including web fonts, input methods, localisation. It will also be a perfect time to ask what you can do to support your language.

We want knowledgeable people to support us by being active in their "language support teams". We want them to verify what we do in MediaWiki and we want people to verify, append and amend what is known about a language in the standards like the CLDR.

Doing this for your language is what allows easy and obvious use of a language. It allows for information that is recognisable as being in your language. Helping you help your language is what we do. Helping you use your language is what we develop in MediaWiki.
Thanks,
     GerardM

Sunday, November 13, 2011

#CLDR, the name of languages in Serbian is in lower case

At #Translatewiki.net we prefer to use information from the CLDR. Sometimes we get a query why a specific message that uses CLDR information is incorrect.
Is there any parser that converts the words between it to lowercase? Value of $1 variable in this message, in Serbian, should be lowercase. So, Translation statistics for Serbian would be Статистика превода за Српски, which won't be gramatically correct since language names in Serbian are written lowercase.
We use the names of languages from the CLDR and we encourage our localisers, our languages support teams to look at the content of the CLDR because we do use it in our MediaWiki software.

When the data as provided by the CLDR is problematic, we can hack it. The fact is that we really really do not like to do that.
Thanks,
      GerardM

Friday, October 21, 2011

#CLDR and #GLIBC - what makes a #standard for locales II

Both GLIBC and CLDR register similar things. They register for similar but different purposes. One is a standard and the other is a de facto standard.

When people with the right background take an interest in a language like Sourashtra, they can have many objectives. What we are looking for are people who are interested in enabling, supporting and promoting the use of their language.

When a font is needed and they can point us to a freely licensed font, we can support them. When a keyboard is needed specifically for a language or a script, we can support them. When they provide us with the code to transliterate from one script to another, we can support them.

To do this for one application is a lot of work and there are so many applications out there, so much replication is happening. Many of these applications are as useful and relevant as MediaWiki. It is for this reason that all the language support teams are asked to verify and possibly amend and append standards like the CLDR or GLIBC.

With information available in the right places and with fonts, keyboard mappings and transliteration scripts shared widely, in a perfect world there will no need to ask to verify, amend and append. By working together using the same data, there will come a time when there is no longer anything to amend or append, when the data for a language new to Wikipedia will just be there.
Thanks,
       GerardM

Monday, September 26, 2011

September 26

In Finland today is the official name day for Kuisma, Finn, Gáivvaš and Johannes, Juhani, Juho. Finland is not the only country with name days. The fun thing is that people can say that it is "Kuisma day" and expect you to know that it is September 26. Well that is to say, when you are Finnish.

When your language is different, the name days are different or they may not exist at all. So what to do. How do you deal when you translate a text that includes name days? Or is it just that way? Should name days be considered as an open standard because they contain public facts or can they be proprietary..

When name days are to be part of a standard, it would make sense for them to be part of a standard like the CLDR. On the other hand, there is so much data that is still missing. Does it make sense to add even more to the CLDR?
Thanks,
      GerardM

Thursday, September 22, 2011

The difference #Urdu makes

When you learn to write, it matters where you learn this skill. You may learn to write in a different script, and the characters may look different from what is considered to be the standard. The examples to the right show fonts that are usable for the Urdu language. Some reflect how Urdu is written manually. for instance in the Nastaleeq style.

Of the examples to the right, the one at the bottom comes with some versions of Microsoft Windows, the others are developed with Urdu in mind by the Center for Research in Urdu Language Processing. As Urdu uses its own set of characters of the Arabic script, it is no wonder that Urdu has its own keyboard mapping as well. This is available from the same organisation..

Urdu is a good example that when you support languages, you cannot take too much for granted. It also indicates that as we develop software intended to support languages we will need people knowledgeable about their language. People who are willing to test the software we develop for all languages.

These same people may help us provide, amend and verify the information that exists in standards like the CLDR for their language. We need to build teams of people willing to support their language. This will not only help us get the best out of MediaWiki, it can have a much bigger impact when we do this well.
Thanks,
     GerardM

Wednesday, August 31, 2011

#Hindi #Wikipedia has 100K or 10.00.00 or one lakh articles

News on the mailing lists has it that the Hindi Wikipedia has one lakh articles. A lakh is 105 and it is one of those numbers we happily recognise as a cause for celebration. So, congratulations to the Hindi community and lets hope that this auspicious occasion will translate in a project that will be increasingly popular.

A lakh is a unit in the Indian numbering system. These numbers are written differently from what most of us are used to. It just happens to be the same as 100.000. Numbering systems are defined in standards and, the appropriate standard is the CLDR.

Given that we do support languages and consequently peculiarities like different numbering systems, there is a Bugzilla bug asking for the support of the Indian numbering system. It has been scheduled for proper attention and, that is a sign of the Wikimedia Foundation doing good for all the cultures it supports with its projects.
Thanks,
       GerardM

Monday, April 04, 2011

The #CLDR does not work for us

A #standard works when it solves problems. #Wikipedia supports more languages than the CLDR and the result is that we cannot rely on the date format being available.

At translatewiki.net we are very much in favour of using data supplied by standards and, we would applaud it when we can work with Unicode to compile data that is its remit. One of the ways could be to have an area with a friendly interface where people can add this information.

The availability of such data could even be considered essential for the running of a new Wikimedia wiki in a language. The benefit would be that we cut the number of requests for changes to values and formats of data in our projects. The benefit for the CLDR will be that for many more languages information will be available.

The CLDR has as part of its process that information is verified by experts. This is also true when a new language gets its first Wikimedia wiki. Asking the expert to take special care for the locale data will help accept the CLDR to accept this data.
Thanks,
      GerardM

Sunday, March 20, 2011

What to do when the #CLDR does not have it II

In #MediaWiki messages we support the plural of items. Its implementation differs per language and at translatewiki.net we use the CLDR as the source for these rules.

This presents its own problem as the CLDR does not do fractions. Saper has just added Bug 28128 - "Consider whether {{PLURAL:}} should handle fractional numbers".

Having proper support for fractions is relevant when you want to inform for instance how many millions the Fundraiser has made for us. The current messages are mostly based on lists so that keeps it nicely integer.

Given the complexity of the support for integers in all our languages, it is a nice can of worms we are opening up. It helps that the Fundraiser only starts in November.
Thanks,
        GerardM

Saturday, March 19, 2011

What to do when the #CLDR does not have it

At #translatewiki.net we use standards. It makes our life a lot easier because it prevents us from arguing about things we do not really know about. When there is a standard, we have always referred to the standard and asked people to update the data in the standard.

The CLDR is a standard where you can enter the name of languages and currencies and time formats... A long list of items is supported. One problem is that the CLDR supports 217 languages. At translatewiki.net we support 344 languages. When our people experience problems with getting data into the CLDR, at some stage we want to reconsider.

As we may be getting into the collection of translations for things like the names of languages, we will do this reluctantly and we will continue to give precedence to the information provided by the standards. Another thing is that there are many lists with the translations of language names and it does not makes sense to do the same thing yet again. One such can be found at OmegaWiki...

Another reason to be reluctant is that working on such functionality keeps our developers from more hard core work needed at translatewiki.net. We are very much like any open source project; there are more then enough ambitions and we welcome all the help we can get.
Thanks,
       GerardM

Monday, February 07, 2011

#Usability at #translatewiki.net

Translatewiki.net has so many changes that there is the option for  "translations only" or "filter translations" on the recent changes. When you combine this with "New messages" providing the messages in threads, you will agree that a lot of thought has gone in getting the message you are looking for..

Another thing that is different is the "language selection". Usually the selection of languages is done by providing the ISO code and the name of the language in that language, but here it is a choice to make the selection more precise. For this reason the names of languages are shown in the language of the user interface.


As you can see of this screenshot of the recent changes in Farsi, many of the languages are still in English and, I have been told that some of the names of languages are wrong. So the question is not only how to add the missing names for languages but also how to improve on what is already there. If memory serves me well, we use the information from the CLDR for this. Translatewiki has a long standing policy to use as much as possible information that is contained in standards so we urge our translators to help out where it has the best effect.

You may also have noticed that the box is not properly right to left as is appropriate for Farsi, but I believe that one is not hard to remedy.
Thanks,
      GerardM