#IPNI is a collaborative project between three august bodies in the taxonomy of plants. They are the Royal Botanic Gardens, Kew, the Harvard University Herbaria, and the Australian National Herbarium.
There are three areas where IPNI sets the standards: plants, authors and publications. The objective is to disambiguate any taxonomic reference to a plant in scientific literature to the correct taxon given the taxon name, its author information, publication information and date.
IPNI publishes several graphs indicating the success of their work. I have been involved in this work as a consequence of a database project I did for my father who loved his cacti and succulents.
One example of what information IPNI provides can be found in this page for the "genus" Echninocactus. In my understanding, the correct full taxonomic name is: "Echinocactus Link & Otto Verh. Vereins Beford. Gartenbaues Konigl. Preuss. Staaten 3: 420. 1827". It has all the required information, it has type information, it has links all as you would expect of a standard like this.
To appreciate the work of IPNI; in stead of "Link & Otto", there may have been: "Link and Otto" or "Link et Otto" or ... obviously the information for the publication is easily made into a different abbreviation.
Wikidata included only a subset of the full taxon information. It is easy enough to understand why; Wikipedia only needed the most current one. It is an easy model; works relatively well and it breaks in the corner cases. With the development of WikiCite there is a great and possibly easy opportunity to expand on the current work given the expanding collaboration with botanical partners like the Biodiversity Heritage Library.
Thanks,
GerardM
Showing posts with label standards. Show all posts
Showing posts with label standards. Show all posts
Saturday, August 05, 2017
Thursday, January 01, 2015
#ISO - the shame that is in #standards behind a #paywall
The International Organisation for Standardisation exists for one reason and one reason alone. It is to define and promote standards. On the one hand it does a wonderful job and on the other hand it has been set up to fail miserably.
Defining a standard is hard work and a political process. Once it has been defined, it is ratified, it has to be adopted. The politicians, that make the process extra hard have for reasons that are ... political, decided that ISO has to pay for itself and consequently it was told to raise funds by charging people for learning about standards.
The consequence is that people make do with other standards or substandard descriptions of the standard like this one for the ISO 8601. Substandard because it is not the standard.
Wikidata includes information about events, it knows about a start date and an end date. This is what the ISO 8601 is about: date, time, calculation of periods and obviously the presentation of all that. Obviously, it is easy enough to buy one copy of the text of the standard. But at the same time, there are many people, projects that are not as fortunate. They are anxious about implementing what is already defined in a standard, an ISO standard.
It is wrong headed politicians that ruin what is important. The politics of standards should be about the macro effects; the adoption of standards. It should not be about micro managing a system and breaking the system in the process.
Thanks,
GerardM
Defining a standard is hard work and a political process. Once it has been defined, it is ratified, it has to be adopted. The politicians, that make the process extra hard have for reasons that are ... political, decided that ISO has to pay for itself and consequently it was told to raise funds by charging people for learning about standards.
The consequence is that people make do with other standards or substandard descriptions of the standard like this one for the ISO 8601. Substandard because it is not the standard.
Wikidata includes information about events, it knows about a start date and an end date. This is what the ISO 8601 is about: date, time, calculation of periods and obviously the presentation of all that. Obviously, it is easy enough to buy one copy of the text of the standard. But at the same time, there are many people, projects that are not as fortunate. They are anxious about implementing what is already defined in a standard, an ISO standard.
It is wrong headed politicians that ruin what is important. The politics of standards should be about the macro effects; the adoption of standards. It should not be about micro managing a system and breaking the system in the process.
Thanks,
GerardM
Saturday, June 30, 2012
#ImpactOCR - Digital publications and the national libraries
According to the #ISBN standard, the "format/means of delivery are irrelevant in deciding whether a product requires an ISBN". However, it is often assumed that a publication requiring an ISBN number is a commercial publication. In the USA and the UK for instance you have to buy your ISBN number or bar code while in Canada they are free because Canada stimulates Canadian culture.
When a standard is not universally applied, it loses application. When all publications are not registered a national library will have to maintain its own system when it is to collect a copy of all publications. As a result the ISBN is dysfunctional as a standard because it does not function as a standard.
When the Wikisourcerers finish the transliteration of a book, it deserves an ISBN number and, national libraries should be aware of these publications. This recognises and registered the work done in the Open Content world. When these books are registered, all the Open Content projects may know that they can concentrate on another book or source.
As the ISBN does not register all publications, it does not do what it is expected to do; function as a standard.
Thanks,
GerardM
When a standard is not universally applied, it loses application. When all publications are not registered a national library will have to maintain its own system when it is to collect a copy of all publications. As a result the ISBN is dysfunctional as a standard because it does not function as a standard.
When the Wikisourcerers finish the transliteration of a book, it deserves an ISBN number and, national libraries should be aware of these publications. This recognises and registered the work done in the Open Content world. When these books are registered, all the Open Content projects may know that they can concentrate on another book or source.
As the ISBN does not register all publications, it does not do what it is expected to do; function as a standard.
Thanks,
GerardM
Thursday, June 28, 2012
#ImpactOCR - A #font using #unicode private use characters
When really old historic texts are digitised and OCR-ed, the images of letters found are mapped to the correct characters. Characters are defined in Unicode and when a character is NOT defined, it is possible to define them in the "private use" space.
As part of the Impact project, really old texts have been digitised, texts in many languages. At the recent presentation it was mentioned by two speakers that there were characters used in the Slovenian and Polish language that are not (yet) defined in Unicode. As part of their project, the missing characters were defined in the Unicode private use area and the scanning software was taught to use them.
With the research completed, with the need for all these characters and their shape defined, it will be great when these characters find their way in Unicode proper. When the code points for the missing characters are defined and agreed, the OCR software can learn to recognise the characters at the new code points, a conversion program can be written for the existing texts and it will be more inviting to include these characters in fonts.
Now that the project is at its end, it is the right moment to extend the Latin script in Unicode even further.
Thanks,
GerardM
As part of the Impact project, really old texts have been digitised, texts in many languages. At the recent presentation it was mentioned by two speakers that there were characters used in the Slovenian and Polish language that are not (yet) defined in Unicode. As part of their project, the missing characters were defined in the Unicode private use area and the scanning software was taught to use them.
With the research completed, with the need for all these characters and their shape defined, it will be great when these characters find their way in Unicode proper. When the code points for the missing characters are defined and agreed, the OCR software can learn to recognise the characters at the new code points, a conversion program can be written for the existing texts and it will be more inviting to include these characters in fonts.
Now that the project is at its end, it is the right moment to extend the Latin script in Unicode even further.
Thanks,
GerardM
Tuesday, May 15, 2012
#CLDR will know language names in #Esperanto
#MediaWiki uses the language names as defined in the CLDR. It is therefore important that people compile a list of the translations for their language make them available to be used as the standard translation.
Arno did exactly that. The list he created contains all the codes as used for Wikipedia combined with the translation in Esperanto. It is an important effort and it would be great when we have such a list for all the other Wikipedia languages as well.
As standards are standards, we were asked to provide the list in an XML format. It took some pastes and find and replaces and it looks good. The only problem is that some of the codes used are not standard codes. Several codes have been removed, "als" for instance is the code for Albanian Tosk not Alleman. There may be some other "language codes" in there that are not recognised in a standard.
When your language can do with additional translations, please follow the Esperanto example and provide a list of translations in your language.
Thanks,
GerardM
Arno did exactly that. The list he created contains all the codes as used for Wikipedia combined with the translation in Esperanto. It is an important effort and it would be great when we have such a list for all the other Wikipedia languages as well.
As standards are standards, we were asked to provide the list in an XML format. It took some pastes and find and replaces and it looks good. The only problem is that some of the codes used are not standard codes. Several codes have been removed, "als" for instance is the code for Albanian Tosk not Alleman. There may be some other "language codes" in there that are not recognised in a standard.
When your language can do with additional translations, please follow the Esperanto example and provide a list of translations in your language.
Thanks,
GerardM
Thursday, May 03, 2012
A #CLDR walkthrough
For #MediaWiki, the CLDR information is important. Sadly for many of the languages supported in a Wikimedia Foundation project the information is not (yet) available. Several things are needed;
So please watch the video and do what you can do for your language.
Thanks,
GerardM
- People who know their language well enough to enter the data
- People who know their language well enough to verify the data entered
These people exist for any language. The question is what does it take for people to enter the data. For the Asturian language, a language from Spain, the data is now being entered.
One of the things that may help is instructional material and, it is quite wonderful that an instruction video has just been released. To quote the message announcing the material:
A new 52-minute walkthrough video is now available, showing how to use the CLDR Survey Tool to enter data, and prioritize your work. The video and explanatory material are available on this CLDR site page.
So please watch the video and do what you can do for your language.
Thanks,
GerardM
Wednesday, April 25, 2012
Supporting #Asturian in the #CLDR II
It is that time when people CAN support their language and enter data about their language in the CLDR. Last year, collecting the necessary data for Asturian did not happen within the set amount of time.
This year another try will be made to find and enter the data for Asturian. It is good news. I hope to learn about more languages and locales that will be entered.
Thanks,
GerardM
This year another try will be made to find and enter the data for Asturian. It is good news. I hope to learn about more languages and locales that will be entered.
Thanks,
GerardM
Sunday, April 15, 2012
Supporting #Asturian in the #CLDR
Finding the data to support a language in the CLDR can be a struggle. The core requirements are only a few so how hard can it be...
- (04) Exemplar sets: main, auxiliary, index, punctuation. [main/xxx.xml]
- (02) Orientation (bidi writing systems only) [main/xxx.xml]
- (01) Plural rules [supplemental/plurals.xml]
- (01) Default content script and region (normally: normally country with largest population using that language, and normal script for that). [supplemental/supplementalMetadata.xml]
- (N) Verify the country data ( i.e. which territories in which the language is spoken enough to create a locale ) [supplemental/supplementalData.xml]
- *(N) Romanization table (non-Latin writing systems only) [spreadsheet, we'll translate into transforms/xxx-en.xml]
When you read this, the text indicating what the initial requirements are, it becomes quite obvious why this process has such a bad reputation.
- It is not clear nor relevant where the data provided ends up in an xml format
- Orientation is very much an aspect of a script, not of a language nor of a locale.
- In the survey tool it is Esperanto that proves that a language may not fit into a locale anyway.
- Romanisation is stated as a requirement. It is however not obvious at all that every script or language has ever been romanised in a standardised way and why this might keep a language out of the standard
In the past there has been an attempt to provide information for the Asturian language to the CLDR. The good news is that there is documentation on why it failed. The problem was that when you establish data about a language, you need to be certain. Four Unicode characters ('Ḥ ḥ Ḷ ḷ') are used for writing the Asturian language properly and the literature on the subject was not consistent. This issue was resolved but it took more time then was available in the CLDR time box.
The Asturian example proves that getting data ready for a standard takes time. The practice of closing a request because the data was not provided within a set amount of time is what stopped people dead in their tracks. We can only hope that Asturians will find what it takes to get support for their language in this time box.
Thanks,
GerardM
The Asturian example proves that getting data ready for a standard takes time. The practice of closing a request because the data was not provided within a set amount of time is what stopped people dead in their tracks. We can only hope that Asturians will find what it takes to get support for their language in this time box.
Thanks,
GerardM
Monday, March 12, 2012
#Standards - A gap in plural support
![]() |
| #multilingweb |
When I wrote about plural, a subject that is actively discussed on translatewiki.net, Niklas told me about his frustration that his inventory plural rules in various databases had not resulted in anything at all.
As the Multilingual Web conference explicitly asks to identify where standards and best practices cover our needs, this is certainly one that is relevant to us.
As the hashtag of the conference will surely find its way to twitter, this can be seen as an experiment; will people who will go to the conference see this and will there be some follow up at the conference.
Thanks,
GerardM
PS Niklas will be at the conference as well
Supporting #plural in #gettext and #MediaWiki
Gettext, the #i18n module of the #GNU software, is and has been really import for the internationalisation and localisation of open and free software. To a large extend it is what is used by many localisation platform.
What gettext provides is technology. What it does not provide is the specific rules needed to implement the internationalisation for a specific language. When we bootstrapped plural support at translatewiki.net, we copied the rules from other applications to start of with.
The way plural is implemented for applications supported at translatewiki.net is well documented. When you read the documentation, it is clear that there is no consistency and these inconsistencies are documented.
One of our contributors, Lloffiwr is taking an active interest in the subject and is compiling a list that shows the MediaWiki plural rules for the languages enabled for localisation at translatewiki.net. Such a list informs our localisers what is expected of them when they localise a message with plural support.
Thanks,
GerardM
What gettext provides is technology. What it does not provide is the specific rules needed to implement the internationalisation for a specific language. When we bootstrapped plural support at translatewiki.net, we copied the rules from other applications to start of with.
The way plural is implemented for applications supported at translatewiki.net is well documented. When you read the documentation, it is clear that there is no consistency and these inconsistencies are documented.
One of our contributors, Lloffiwr is taking an active interest in the subject and is compiling a list that shows the MediaWiki plural rules for the languages enabled for localisation at translatewiki.net. Such a list informs our localisers what is expected of them when they localise a message with plural support.
Thanks,
GerardM
Wednesday, January 25, 2012
A doctor a day keeps the #Apple away
The #Wikimedia Foundation is into education. Apple is in the business of making money, preferably to the exclusion of others. Recently it proposed "new educational" tools that were supposed to make for a rich educational experience.
There is nothing wrong with providing tools that will provide students with a rich educational experience. However, when it breaks the standards that allow for the cooperation on such an experience, combine this with a EULA intended to prevent further distribution it proves rotten to the core.
For me it is clear, Apple moved completely and utterly to the dark side. The benefit of its possible superior technology is completely offset by the negative impact it has on the rest of society. Apple used to compete and do well because of its superior products and its customers were willing to pay a premium for this. By excluding cooperation and interoperability, the Wikimedia Foundation cannot sign up to what have the potential of becoming a great educational product. It is sad that it has been poisoned by unacceptable conditions.
Thanks,
GerardM
There is nothing wrong with providing tools that will provide students with a rich educational experience. However, when it breaks the standards that allow for the cooperation on such an experience, combine this with a EULA intended to prevent further distribution it proves rotten to the core.
For me it is clear, Apple moved completely and utterly to the dark side. The benefit of its possible superior technology is completely offset by the negative impact it has on the rest of society. Apple used to compete and do well because of its superior products and its customers were willing to pay a premium for this. By excluding cooperation and interoperability, the Wikimedia Foundation cannot sign up to what have the potential of becoming a great educational product. It is sad that it has been poisoned by unacceptable conditions.
Thanks,
GerardM
Saturday, October 01, 2011
The use of outside #standards
The notion that we are going to filter images maybe articles from our users makes me feel uncomfortable. The amount of bad faith in the discussion makes me feel nauseous; I have stopped following the multiple threads on the subject. The arguments why a single set of arguments should not be used are in my opinion compelling. The arguments why we do not want to set the values that enable the loss of immediate visibility are equally compelling.
Consequently I would like to steal a page out of the language policy; have an external body with sufficient authority set for us us up with the necessary values THEY need to do their job. We in turn can piggy back and enable the values set and allow our readers to select one of the levels provided they feel comfortable with.
So here is the deal; The IEEE LOM is a standard used by many educational organisations to mark electronic information that can be used on demand by students. The information is tagged in many ways including an age level that corresponds with what a teacher consider to be appropriate for a student.
There are multiple benefits to this scheme. It makes our content discoverable in an automated way from within education. Our content is rated for its usability and this provides us with feedback on our content. I am sure that many of our articles will be rated as "too long did not read" or using a vocabulary that is too demanding and consequently get a high age rating as a consequence.
Given that a lot of money is spend on providing quality educational information by many educational entities, they will be more then happy to do the rating. When you combine this with the use of FlaggedRevisions to store the information it seems a good fit for many reasons.
And yes, we can ask for contributions for hosting the data and providing content on a just in time basis.
Thanks,
GerardM
Consequently I would like to steal a page out of the language policy; have an external body with sufficient authority set for us us up with the necessary values THEY need to do their job. We in turn can piggy back and enable the values set and allow our readers to select one of the levels provided they feel comfortable with.
So here is the deal; The IEEE LOM is a standard used by many educational organisations to mark electronic information that can be used on demand by students. The information is tagged in many ways including an age level that corresponds with what a teacher consider to be appropriate for a student.
There are multiple benefits to this scheme. It makes our content discoverable in an automated way from within education. Our content is rated for its usability and this provides us with feedback on our content. I am sure that many of our articles will be rated as "too long did not read" or using a vocabulary that is too demanding and consequently get a high age rating as a consequence.
Given that a lot of money is spend on providing quality educational information by many educational entities, they will be more then happy to do the rating. When you combine this with the use of FlaggedRevisions to store the information it seems a good fit for many reasons.
And yes, we can ask for contributions for hosting the data and providing content on a just in time basis.
Thanks,
GerardM
Monday, September 26, 2011
September 26
In Finland today is the official name day for Kuisma, Finn, Gáivvaš and Johannes, Juhani, Juho. Finland is not the only country with name days. The fun thing is that people can say that it is "Kuisma day" and expect you to know that it is September 26. Well that is to say, when you are Finnish.
When your language is different, the name days are different or they may not exist at all. So what to do. How do you deal when you translate a text that includes name days? Or is it just that way? Should name days be considered as an open standard because they contain public facts or can they be proprietary..
When name days are to be part of a standard, it would make sense for them to be part of a standard like the CLDR. On the other hand, there is so much data that is still missing. Does it make sense to add even more to the CLDR?
Thanks,
GerardM
When your language is different, the name days are different or they may not exist at all. So what to do. How do you deal when you translate a text that includes name days? Or is it just that way? Should name days be considered as an open standard because they contain public facts or can they be proprietary..
When name days are to be part of a standard, it would make sense for them to be part of a standard like the CLDR. On the other hand, there is so much data that is still missing. Does it make sense to add even more to the CLDR?
Thanks,
GerardM
Thursday, September 22, 2011
The difference #Urdu makes
When you learn to write, it matters where you learn this skill. You may learn to write in a different script, and the characters may look different from what is considered to be the standard. The examples to the right show fonts that are usable for the Urdu language. Some reflect how Urdu is written manually. for instance in the Nastaleeq style.
Of the examples to the right, the one at the bottom comes with some versions of Microsoft Windows, the others are developed with Urdu in mind by the Center for Research in Urdu Language Processing. As Urdu uses its own set of characters of the Arabic script, it is no wonder that Urdu has its own keyboard mapping as well. This is available from the same organisation..
Urdu is a good example that when you support languages, you cannot take too much for granted. It also indicates that as we develop software intended to support languages we will need people knowledgeable about their language. People who are willing to test the software we develop for all languages.
These same people may help us provide, amend and verify the information that exists in standards like the CLDR for their language. We need to build teams of people willing to support their language. This will not only help us get the best out of MediaWiki, it can have a much bigger impact when we do this well.
Thanks,
GerardM
Of the examples to the right, the one at the bottom comes with some versions of Microsoft Windows, the others are developed with Urdu in mind by the Center for Research in Urdu Language Processing. As Urdu uses its own set of characters of the Arabic script, it is no wonder that Urdu has its own keyboard mapping as well. This is available from the same organisation..
Urdu is a good example that when you support languages, you cannot take too much for granted. It also indicates that as we develop software intended to support languages we will need people knowledgeable about their language. People who are willing to test the software we develop for all languages.
These same people may help us provide, amend and verify the information that exists in standards like the CLDR for their language. We need to build teams of people willing to support their language. This will not only help us get the best out of MediaWiki, it can have a much bigger impact when we do this well.
Thanks,
GerardM
Friday, September 16, 2011
The name is Bon, James Bon II
Santhosh started work on an #Unicode UAX31 implementation.. Good for us because we support all the scripts for all the languages we have projects in. Our best practice has it that people can have a user in their own language.
The UAX31 standard defines the standard rules for our best practice. Writing a reference implementation that will actually be used is a bit different; the Wikimedia Foundation implemented a black list of names for instance.
Once you start writing an implementation, you encounter all the ambiguities, all the issues that are still open. How for instance do you cope with the Arabic script, what to do with a "." or a full stop in the middle of a user?
It makes sense to allow for the characters that are used in a language. This implies that knowing what language to expect is crucial. There are two obvious approaches; you expect a new user in the default language or the languages is defined in the preferences.
When you are interested in this subject have a read of the current version of the code.
Thanks,
GerardM
The UAX31 standard defines the standard rules for our best practice. Writing a reference implementation that will actually be used is a bit different; the Wikimedia Foundation implemented a black list of names for instance.
Once you start writing an implementation, you encounter all the ambiguities, all the issues that are still open. How for instance do you cope with the Arabic script, what to do with a "." or a full stop in the middle of a user?
It makes sense to allow for the characters that are used in a language. This implies that knowing what language to expect is crucial. There are two obvious approaches; you expect a new user in the default language or the languages is defined in the preferences.
When you are interested in this subject have a read of the current version of the code.
Thanks,
GerardM
Thursday, September 08, 2011
The name is Bon, James Bon
Santhosh is one of the special agents in the fight to bring language support to the Internet. Identifying him in English is easy; all the characters used to transliterate സന്à´¤ോà´·് à´¤ോà´Ÿ്à´Ÿിà´™്ങല് are available to identify Santhosh for who he is.
The last character in his name is the "zwj". According to some, this character is not available for identification purposes. Without the "zwj", the name looks different:
It becomes more interesting when you write Sri Lankha in Singhala. this cannot be done without a "zwj". From a Wikimedia Foundation point of view, the Unicode report "Unicode Identifier and Pattern Syntax" assumes for many languages that they are "aspirational" or "limited" use, is not really workable. Our aim is to have support for all scripts and identifying people by their name; their real name.
As we do identify people, an implementation of this Unicode specification is important to us. Having people like Mr à´¤ോà´Ÿ്à´Ÿിà´™്ങല് in the drivers seat will surely get us a best result. It may even get us a reference implementation.
Thanks,
GerardM
The last character in his name is the "zwj". According to some, this character is not available for identification purposes. Without the "zwj", the name looks different:
- സന്à´¤ോà´·് à´¤ോà´Ÿ്à´Ÿിà´™്ങല്
- സന്à´¤ോà´·് à´¤ോà´Ÿ്à´Ÿിà´™്ങല്
It becomes more interesting when you write Sri Lankha in Singhala. this cannot be done without a "zwj". From a Wikimedia Foundation point of view, the Unicode report "Unicode Identifier and Pattern Syntax" assumes for many languages that they are "aspirational" or "limited" use, is not really workable. Our aim is to have support for all scripts and identifying people by their name; their real name.
As we do identify people, an implementation of this Unicode specification is important to us. Having people like Mr à´¤ോà´Ÿ്à´Ÿിà´™്ങല് in the drivers seat will surely get us a best result. It may even get us a reference implementation.
Thanks,
GerardM
Monday, July 25, 2011
#Unicode is not that complicated
When good people associate Unicode with this #XKCD strip, things are wrong, seriously wrong. The value of a standard is in its acceptance. Unicode is best known for its work on scripts and it is this work that makes Wikipedia possible. Without the standardisation brought about by Unicode it would be impossible to support the 270+ languages who have a Wikipedia and, the languages working towards their Wiki in the Incubator. As there is confusion about Unicode lets analyse what is said. First of all, Unicode is a work in progress. As a consequence there is no guarantee that every script has been encoded and, there is even less of a guarantee that a font is available let alone a freely licensed font. One reason why the situation is not good is because many organisations consider the development of Unicode encodings for scripts and fonts out of scope.
Unicode is developing an additional standard, the CLDR or the Common Locale Data Repository. While this standard is important, many of the languages supported by the Wikimedia Foundation are not represented and languages that are represented do not have complete data. Consequently there is no authoritative source for many a date format, a currency format, the sorting order ...
Unicode is a consortium of companies and organisations. Particularly the companies have an interest in ensuring that the technology represents its investments. When individuals are singled out negatively, particularly people like Michael Everson who does not represent company interests, who is involved in open content, it becomes clear that Unicode should to do what it is supposed to do and provide support for languages, all languages and scripts, all scripts.
Thanks,
GerardM
Tuesday, May 17, 2011
#mwhack11; Managing the #stack of #MediaWiki
#LAMP is considered to be that stack. It could also be seen as a pyramid of functionality. When you look at this illustration for LAMP is that "Application" represents any application.
As MediaWiki is the application that runs Wikipedia, it is obvious that language related technology in the LAMP stack is of special relevance to us. With some 270 Wikipedias in as many languages it has a coverage that is bigger then some of the standards that should cover "our" languages as well as all the others.
At translatewiki.net we feel the pain of the missing data. We feel it for MediaWiki, for Mifos, for many of our applications. To alleviate this pain, we can use what we have. We know what is in the CLDR, LIB-C; this can be merged. We have a community representing all the languages we support who we can ask the data for their language.
We have discussed this in the WMF language committee at the Berlin Hackathon, we discussed it with the translatewiki.net managers. It is now a matter of finding out how best to do it, do it and get the data that is new into the standards where it belongs if it to support the whole stack.
Thanks,
GerardM
As MediaWiki is the application that runs Wikipedia, it is obvious that language related technology in the LAMP stack is of special relevance to us. With some 270 Wikipedias in as many languages it has a coverage that is bigger then some of the standards that should cover "our" languages as well as all the others.
At translatewiki.net we feel the pain of the missing data. We feel it for MediaWiki, for Mifos, for many of our applications. To alleviate this pain, we can use what we have. We know what is in the CLDR, LIB-C; this can be merged. We have a community representing all the languages we support who we can ask the data for their language.
We have discussed this in the WMF language committee at the Berlin Hackathon, we discussed it with the translatewiki.net managers. It is now a matter of finding out how best to do it, do it and get the data that is new into the standards where it belongs if it to support the whole stack.
Thanks,
GerardM
Thursday, April 14, 2011
#Lisa is insolvent - #Standards come at a cost
Lisa is an organisation involved in standards for the translation industry. In order to get an industry to accept standards, there is a need for an organisation that not only helps formulate standards but also promotes their use. Lisa organised worldwide conferences where standards and best practices were presented, promoted and defined.
With its insolvency an organisation that played its role in the definition of standards like TMX and TBX. Let's hope that they find a way to continue the good work.
Thanks,
GerardM
With its insolvency an organisation that played its role in the definition of standards like TMX and TBX. Let's hope that they find a way to continue the good work.
Thanks,
GerardM
Thursday, April 07, 2011
www.wikipedia.org is not available for Wawa
The #URL for a #Wikipedia consists of an ISO-639 code followed by "wikipedia.org". Because of the configuration of the Wikipedia domain, "www" is not necessary in front of an URL.
The Wawa language of Cameroon has as its ISO-639 code "www" and, this poses a problem; when you type www.wikipedia.org, you go to the portal. A functional URL for Wawa would be www.www.wikipedia.org.
This is quite cumbersome and this has been recognised as one of those corner cases that pose a problem. As we have someone from SIL on the language committee, the request for a Wawa Wikipedia came to the attention of Melinda Lyon. She has now proposed to change the code for Wawa to "wwx".
Thanks,
GerardM
The Wawa language of Cameroon has as its ISO-639 code "www" and, this poses a problem; when you type www.wikipedia.org, you go to the portal. A functional URL for Wawa would be www.www.wikipedia.org.
This is quite cumbersome and this has been recognised as one of those corner cases that pose a problem. As we have someone from SIL on the language committee, the request for a Wawa Wikipedia came to the attention of Melinda Lyon. She has now proposed to change the code for Wawa to "wwx".
Thanks,
GerardM
Subscribe to:
Posts (Atom)



















