Showing posts with label transliteration. Show all posts
Showing posts with label transliteration. Show all posts

Monday, June 16, 2014

#Wikidata - କବି ପ୍ରସାଦ ମିଶ୍ର

In my ongoing project to document the deaths of notable people in 2014, a friend helped me with Mr କବି ପ୍ରସାଦ ମିଶ୍ର. He knows Odia, a language that Google translate cannot help me with.

It was obvious that when you only know someone as Mr କବି ପ୍ରସାଦ ମିଶ୍ର, nobody will find him if they do not know that this can be transliterated to Kabi Prasad Mishra. As this transliteration was added, we can now Google for this gentleman and find more information about him.

Subhashish asked my help in turn for a transliteration for Mr Герич Ігор Дионізович. This person is Ukranian and to be honest, I do not know how to transliterate his name into English.

Plenty of opportunities left in Wikidata :)
Thanks,
     GerardM

Tuesday, November 19, 2013

The Case for Localizing Names, "part 3"

My friend Amir wrote twice ([1], [2]) about the need for the localisation of names. The need for localisation is rather obvious when the original name is in another script. But names written in the Latin script can be as foreign. Take for instance the Czech author and screenwriter Jiří Růžička, I do not know how to type his name and, I am sure 98% of the users of Wikidata will be able to do so either.

When a name uses characters that are not in use in a language, it follows that the name is unlikely to be found in a language and consequently transliteration is in order. The name as originally written is not an alias. It is something different. It probably needs an attribute of the type "multilingual text" for support in Wikidata.
Thanks,
       GerardM

Saturday, June 30, 2012

#ImpactOCR - Digital publications and the national libraries

According to the #ISBN standard, the "format/means of delivery are irrelevant in deciding whether a product requires an ISBN". However, it is often assumed that a publication requiring an ISBN number is a commercial publication. In the USA and the UK for instance you have to buy your ISBN number or bar code while in Canada they are free because Canada stimulates Canadian culture.

When a standard is not universally applied, it loses application. When all publications are not registered a national library will have to maintain its own system when it is to collect a copy of all publications. As a result the ISBN is dysfunctional as a standard because it does not function as a standard.

When the Wikisourcerers finish the transliteration of a book, it deserves an ISBN number and, national libraries should be aware of these publications. This recognises and registered the work done in the Open Content world. When these books are registered, all the Open Content projects may know that they can concentrate on another book or source.

As the ISBN does not register all publications, it does not do what it is expected to do; function as a standard. 
Thanks,
       GerardM

Wednesday, June 13, 2012

Playing the Wikisourcerer

For any #Wikisource, the proofread page extension provides a must have functionality. Given that it is used on every Wikisource, it is reasonable to expect consistent functionality. The screenshot below shows the the proofreading page for "Noodlot" a book by Louis Couperus. As you can see, the page numbers are in blue or red and they show the classic MediaWiki behaviour. When you compare this with any index page for proofreading on the English language Wikisource, you will find the numbers in multiple colours indicating its position in the proofreading work flow.


As Wikisource is primarily a workflow environment, it is crucial to have a complete implementation of the tooling. When asked, it was indicated that many of the small fixes happen exclusively on the English Wikisource. This lack of support for the proofreading extension elsewhere is an additional argument for doing away with all the single language Wikisources. When one Wikisource provides adequate tooling for the workflow, another wiki can be used to publish its finished content.

A Wiki publishing content in a final form can publish on behalf of any and all open content project that create finished products. Such an expanded project where finished content is marketed to our public will achieve multiple goals:
Thanks,
      GerardM

Friday, March 09, 2012

Transliterate when it grows your audience

The Chinese and the Serbian Wikipedia have one thing in common; their content is shown in one of the two scripts you can select from. In essence the process is simple; you change a text from one to another script using a fixed set of rules and the only difference is the script of the language. You do not change the orthography it should be just other characters saying the same thing.

Understanding this is quite important because transliteration is not to accommodate differences in dialects. Far from it. When dialects are expressed in the same script, the differences become easier to understand when there is no longer any confusion because of the different scripts.

The "InScript" input methods exist for the scripts of India. What makes them special is that the same sounds are placed at the same location. This makes typing in different scripts easy. This gives the impression that it should be relatively easy to transliterate between scripts.

Changing scripts for a text is of relevance for Sanskrit. Sanskrit is written in many scripts and when a text originated in a script different from Devanagari, many readers of such an original text are helped with a transliteration into a script they are familiar with. At the same time it helps people appreciate how broad a cultural base the Sanskrit language has.

When transliteration works for Sanskrit, it is likely that the same or similar routines will work well for languages like Konkani. At Silpa there is a tool where you can test transliteration. It is a work in progress; each script has many features that need attention.Custom logic need to be written sometimes for script pairs, sometimes for specific language attributes.

What would be cool is when someone works on this existing code for the transliteration of Indian scripts. It needs more work both on script specific rules and on languages specific rules.
Thanks,
     GerardM

Saturday, October 01, 2011

What #script for #Konkani

When a language is written in multiple scripts, five scripts to be exactly, writing about each subject five times is not really an option. Even English with only one script is not done writing down the sum of all knowledge.

When it is theoretically possible to transliterate from and to multiple scripts, things start to look up. The trick for Konkani will be how to support all scripts including the Arabic script.

It is normal to drop the vowels when using the Arabic script. This makes transliterating to the Arabic script possible but it is impossible to transliterate to any of the other scripts.

The question is, would it be ok to force Konkani people to include the vowels and is  this enough to make transliteration from the Arabic script possible?
Thanks,
       GerardM

Wednesday, August 03, 2011

Transcription of Ladino


At the #Wikimania hacking days, in a corner, Can Evrensel was happily working on a project that intends to transcribe Ladino from the Latin script to the Hebrew script and transliterate Ladino to the different orthographies it is known for.


As Ladino is written in different ways, it is wonderful to learn that research is done to see if it is possible to represent Ladino in ways that are familiar to the people who know the language. Research like this invigorates a language, it gives it a lease of life.
Thanks,
      GerardM

Monday, March 14, 2011

A #Wikipedia in the #Tagalog in the Baybayin script?

Hell no. The rules are clear; one project for one language. So without much fanfare the request was denied.

This is not the end of the story.

There is a solution open to be able to read the Tagalog Wikipedia in the Baybayin script. There is even room for a user interface in Baybayin Tagalog. The solution comes in the form of transliterating the existing Latin text.

Transliterating Filipino is not exactly new; there is a website where Tagalog is transliterated in two distinct ways; in the original Baybayin script and in the Baybayin script with Spanish modifications.


If there is one problem, it will be that there is no Unicode font for it. There is a solution for that and all it takes is effort and or money.
Thanks,
      GerardM