Showing posts with label Wiktionary. Show all posts
Showing posts with label Wiktionary. Show all posts

Friday, October 18, 2013

#Wikidata needs 3,780,000,000 labels

For an item to be found in Wikidata, it needs a label. Ideally any item has a label in every language that Wikidata supports and currently there are 280+ languages. Currently Wikidata is not really useful for practically all of these languages. The statistics show that a small percentage has more than 10 labels and most labels have only one label.

There are over 13.5 million items and consequently there is a need for 3,780,000,000 labels give or take a few.

There are several ways to drastically improve the number of labels. There are also ways to minimise the impact of missing labels.

The most obvious one is to use every title of a Wikipedia article as a label. There are for instance 729,712 articles in Chinese and 244,699 articles in Arabic these will probably provide the most requested information in those languages. We can also use lexical information from any freely available resource. Wiktionary for instance includes a wealth of usable data. Implementing these two strategies alone will have a big impact.

The names of people typically stay the same in most languages with the same script. There are standards for transliteration and they can will us with an adequate result. Adding these strategies to the mix and it will become even better.

Finally, there is our community. When we provide them with awareness what labels are most often requested and failed they may either add a label to an existing item or they can add a new item.

Statistics indicate that we have a problem but with this awareness we can build statistics that will indicate what works and where our efforts have the most impact in making Wikidata truly useful.
Thanks,
       Gerard

Wednesday, June 06, 2012

Publishing final versions

On #Wikipedia, an article never finds its final form. There is no final form, there is always a potential next iteration, never mind the quality of the current version. The same is true for Wiktionary; there is always room for another translation, another attribute.

When you want final forms, you produce something static. There are many reasons why this is done. One reason is to create a final version and create a collection of articles and calling it a book. Such books can be priced and it is quite legitimate to sell them on ebay. This has been done and when another such publication is found there are usually some people who complain about it.

For Wikisource and Wikibooks however there are finished products. In Wikisource a project can be considered finished once the proof reading of a digital text has been completed of what started as a scanned text. Only when it is finished, it is ready for easy public consumption.

What Wikisource does is provide the tools for this process. Essentially Wikisource serves a community that is engaged in a workflow. What it does not do so well is publish the finished products and find an audience for them.

The work at Wikisource is important. Its finished products are relevant and they are worthwhile. They fit in our aim to share in the sum of all knowledge. We can give things away, we can ask for a contribution but we owe it to our public that we make a best effort in making all this knowledge we are the custodian of easily accessible.

We may need a new project .. "Wikipublish" where we publish final versions of our "Wikisources". Where we publish collections of Wikipedia articles. Where we publish our "Wikibooks". We do not need much marketing, we need to make our finished goods available to a public and provide it in a format that it finds easy to digest.
Thanks,
       GerardM

Saturday, April 28, 2012

#Wikidata may support what #Wikipedia does not know

The first actual Wikidata project starts to become functional. It is interwiki links on steroids. Just one database for all the interwiki links. An API or application programming interface to acquire the relevant data. At this time it works alongside the old interwiki system and it is getting most of its updates that way.

Work is under way to make an editing interface. when this is fully functional and when Wikidata can cope with the amount of data Wikipedia will ask it to serve it will replace the existing interwikis and good riddance.

Wikipedia is not the only project that makes use of interwiki links. Wiktionary is another. There are fewer Wiktionaries than Wikipedias but languages are treated differently; the English Wiktionary alone supports entries in some 450 languages.


Now consider what happens when the links from Wiktionary and Wikipedia are joined and the translations for concepts in Wiktionary are available in Wikidata ...

Over 150 languages are added. That is exciting enough. More useful will be all the translations to languages that we already support. When a search item is entered, it can be found as it is known to Wikidata and it can be shown in red for easy editing. As the concept found is associated with existing Wikipedia articles, an article can be presented in another language.

How cool is that?
Thanks,
      GerardM

Tuesday, April 03, 2012

#Chennai hackathon IV

Several of the Chennai hackathon projects were interesting because of the different approach they took to known problems. A hackathon is a perfect place to experiment; there are many people around and they often have their own ideas about how things should be done.

Wikiquotes via SMS
You send the name of a person and in reply you get quotes from that person. The idea makes sense but the problem is that the quotes in Wikiquote are not available in a structured way. When quotes are individually tagged, an application to send SMS messages becomes possible from the source. The approach of this project was to copy quotes into a database and use this.

When the Wikiquote communities start using templates for each quote, it will become possible to use Wikiquote itself.

Translation of Gadgets/UserScripts to tawiki
There is no doubt, the translation or localisation of software makes a hell of a difference to the usability of software. The standard approach is to localise at translatewiki.net but gadgets and user scripts are not yet supported in this way.
It is wonderful to learn that the ProvIt and the TwoColumn gadget are now available in Tamil. The challenge left is how to make them available in other languages and how to maintain them when the code needs changing.

Lightweight offline Wiki reader
Even though there are several off line readers like OkaWix and Kiwix, it was the considered opinion that there is room for another one, a reader that can be used when there is only a little room available. The Qvido project was revived at this hackathon and is now available with "build" instructions.

Open source projects exist as long as there is an interest in developing a project and as long as there are people who benefit from it. Being able to host MediaWiki content off line as widely as possible is certainly reason enough for this project to get attention.

Program to help record pronunciations for words in tawikt
When you learn a foreign language it is really hard to get the pronunciation of words right. A program that allows the recording of some 500 words in half an hour is really useful. However, getting it uploaded to Commons is equally important because this can be an enormous time-sink and this is one part of the puzzle that still needs solving.


WikiPronouncer
Another approach to recording pronunciations was to do this on an Android phone. Developing an app in a day is too much of a challenge.. What these two projects make clear is how much pronunciations is a feature that is very much in demand.

As soundfiles are used extensively on other wiktionaries, it does make sense to learn the best practices elsewhere. The least they need is a Tamil makeover eh, a localisation.
Thanks,
       Gerard

Wednesday, February 29, 2012

DYK #Wiktionary does #mobile

To my complete surprise, the "other" Wikimedia projects are now supported by the Wikipedia Mobile software. It is a bit of a hack because you may get a security message on some browsers stating "wikipedia" as the target with you having requesting your project in stead.

Having a menu on the English Wiktionary featuring "unflappable" as the word of the day is appropriate.

When there is no main page for your Wiki yet, you want to read the instructions on the creation of a mobile main page here. The tags available are very much appropriate for a Wikipedia but they can be used to suit your needs. Obviously for the best result you have to complete the localisation of the mobile application at translatewiki.net.
Thanks,
      GerardM

Sunday, October 02, 2011

Looking forward to write #agile user stories

A good joke has a kernel of truth. Agile allows for less documentation, maybe even no documentation and a developer does not need to plan beyond the duration of a sprint. When you start with agile it first goes all haywire and yes the writing of code is what delivers functionality.

I will not be writing code, I will be writing about the functionality that is being developed or has been delivered. Such stories will be in English, short and sweet. They may contain screen shots or pictures that amuse and they are intended to have more people find their way in the most multi-lingual software of all; MediaWiki.

Functionality developed with agile will not appear fully formed like Pallas Athena out of the head of Zeus. It will grow more organically. As experience shows users will surely find ways to use the delivered functionality in ways that were not considered.

As we learn what works and why or what does not work and why, we will collect your stories. These stories are in turn converted in agile user stories and may be decomposed in multiple related stories. Such stories end up in the product backlog and will be prioritised and maybe developed.

There is for instance a user story in this example of Wiktionary template hell:
* Japanese: {{t|ja|スクラム|tr=sukuramu|sc=Jpan}}
The  story may be about editing Japanese, it can be about how the Japanese text was entered or how it has to be shown to a reader. We do want to learn from you what this is all about. For this we need you and becoming part of a language support team is a short cut to getting our attention.
Thanks,
        GerardM

Thursday, July 07, 2011

Digitizing a #Malayalam #dictionary

Digitizing a Malayalam book is quite different from digitizing an English book. For an English book you use OCR and move straight on to the proofreading and formatting of the text.


It becomes even more interesting when the text is rich in all kinds of annotations like in this dictionary. It is the the first Malayalam-English-Malayalam Dictionary by Dr Herman Gundert. There are all kinds of opportunities here, ehm problems.

There is no OCR for Malayalam yet so it has to be crowd sourced. The annotations are cryptic and make only sense in a dead tree dictionary. When you digitize such a text and it remains a flat text and consequently it is extremely hard to use the public domain content elsewhere; in OmegaWiki or in Wiktionary for instance.


Santhosh is experimenting with Semantic MediaWiki . It will allow for exportable information and, that will make the data gained much more useful.
Thanks,
       GerardM

Sunday, April 24, 2011

#Myanmar #Wiktionary reaches 100,000 words!


I started contributing to Wiktionary at the end of last year. It was difficult to maintain a nearly inactive Wiktionary with a couple of regular members. So I decided to make the Wiktionary usable to certain extents first while I tried to recruit new contributors. I collected opensourced dictionary data online such as Ornagai, electrical engineering dictionary porject and Sealang Burmese dictionary. Online community generated data such as Ornagai has a lot of spelling mistakes in it and the format and definitions are inconsistent which need to be improved greatly. Then Myanmar NLP kindly provided me with the Myanmar lexicon database.

Pywikipediabot is a very handy and useful application. I used to run it for fixing minor encoding mistakes in Burmese. Now I used it to upload dictionary data. I spent the Myanmar new year (13th-17th April) formatting the data to Wiki markup. Now Myanmar Wiktionary is available in two languages as English-Burmese and Burmese-Burmese. While bot was uploading data, we tried to improve the homepage design and to add more words which were not in the collected dictionary data. Today, Myanmar Wiktionary reaches 100,000 words.

Unicode Font Usability

It might be confusing that why there are not a lot of Burmese Wikipedians or Wiktionarians. It was a complex drama that there were several pseudo Unicode fonts (they use Unicode codepoints for Burmese, but never follow the codepoints exactly or the encoding order) when standard Unicode fonts were still in development, and one of the pseudo Unicode font named Zawgyi became popular. Unicode fonts are used in Government offices and international projects such as Wikipedia. Yet a normal Burmese online citizen would only use Zawgyi font. We, the Unicode activists, tried to make awareness of Unicode standard and got some achievements as popular IT forums such as MyanmarITPros, Mystery Zillion and Mmitd between 2011 English new year and Burmese new year.


Before major OS vendors such as Apple and Microsoft support Burmese language in OSX and Windows, there should be solution for usability of Wikimedia projects for OS unsupported languages.

1. Font embedding


Burmese script is  one of many branches of Brahmi script. It needs complex shaping and reordering of glyphs. So Unicode fonts with proper contextual rendering are necessary. Most of the people don't have Burmese font or have Burmese font but Zawgyi. They wouldn't be able to see the Wikipedia or Wiktionary texts or wouldn't see them properly. So embedded font is needed for reading purpose. There are several opensourced fonts for Burmese in multiple platforms such as Windows, Linux, OSX and iOS. Some of them are cross-OS enabled. OS and browser detection javascripts were developed and TTF, compressed TTF, eot and woff can be embedded accordingly. As I am not a coder, I would be very glad if someone were to develop a font embedding module for MediaWiki and implement it in local wikis.

2. Build-in keyboard

Even after texts are readable, editing texts in local languages needs special keyboard inputs. Languages such as Burmese needs keyboards with reordering capability. People usually don't have keyboard installed in their computer or mobile devices. There is a project called Narayam which is a build-in keyboard plugin for MediaWiki and I helped writing Myanmar Unicode keyboard for it. There is also a Burmese local project called Keymagic and one of it's effort was web keyboard. I hope with the help of them, editing Wikipedia even in local language would be possible.

@=={Lionslayer>

Wednesday, March 30, 2011

Using the Ş or Ș In #Romanian

The definition of the subset of #Unicode characters used for the Romanian language is quite clear; only the Ș and the ș is correct for the șe. This does not mean however that everybody uses the comma under the S and not a cedilla.

Before Unicode became common place typing the comma under a character was really hard. As a consequence many, many expressions of Romanian are erroneous. People got used to writing with a cedilla.

With the later versions of MS Windows, the keyboard mapping for Romanian made it easy to write correct Romanian. For the people stuck with a wrong keyboard we can easily have Narayam provide a modern input method for Romanian.


This is currently a big deal for the Romanian and English language Wiktionaries. They are in the process of correcting every occurrence of a wrongly written șe and țe. This is quite an undertaking because it affects interwiki links to other Wiktionaries as well.

It also means that they want to ensure that only a correct șe or țe is written in Romanian. This is complicated by the fact that a Ş is correct in for instance Turkish. Being able to identify a text for its language is therefore quite important.

The solution currently implemented on the Romanian Wikipedia is that any t or s with a cedilla is converted to a proper șe or țe. As a consequence Turkish names of people and places are likely to be spelled incorrectly.
Thanks,
      GerardM

PS Please note that the font used for the title does not cope with these characters.

Sunday, March 27, 2011

Supporting multiple languages in one article

Many #Wikipedia articles include text in other languages. There are many reasons why this is the case but technically these foreign bodies are not marked as such. For people there is no problem, it is easy to spot what is in a different language. For computers rendering a text is not much of a problem either; if a character is not available in one font it will likely be in another.

When it is reasonably certain that the majority of our audience do not have a font, current practice is to include a screen shot. While functional, it is not the best solution because search engines will not be able to read these. For Wikimedia projects like Wiktionary, including foreign text is the rule and not the exception and many of its readers see the rectangles of the Unicode fonts missing on their system.

With multiple languages present in a text, a spell checker will mark much of a text as in error. It does this because it is not aware of changes in language.

With the languages in our articles properly marked, a spell checker can prevent obvious errors, a search engine can do its job and we can provide additional targeted support by providing input methods and web fonts.
Thanks,
      GerardM

Sunday, February 21, 2010

Mobile #Wiktionary

I blogged about mobile #Wikipedia and e-mailed with the developer and I learned several new things. The most relevant is that you can create a mobile main page. It then takes some arcane arts to configure this... it is not clear where to look for an how to.

Wiktionary is likely to benefit even more from a mobile interface. The type of information lends itself easily to be presented on a mobile phone. Given the process, it should even be easy to configure this.

I did update the information on translatewiki for Mobile MediaWiki.. With a reference to the "how to" for a configuration, it will be possible to make any MediaWiki installation mobile. That would make MediaWiki a lot more attractive.
Thanks,
      GerardM

Thursday, August 06, 2009

Apple does not like Wiktionary

Apple computer rejected an application for its telephone. The fact that they have funny notions about what its customers can buy is nothing new. What is of interest to me is that the Ninjawords dictionary is based on Wiktionary.

The argument used is that Ninjawords provides access to "other more vulgar words". When people ask me about vulgar words, I always argue that it has a clear benefit when you can learn what has been said to you. Particularly when it is also enriched with etymological information... I know that many people do not really know what they say.
Thanks,
      GerardM

Friday, January 02, 2009

Portuguese

The BBC informed me that the lusophone countries have started a process to adopt new spelling rules. These spelling rules are especially welcomed by countries like Angola who struggle to raise literacy in their country; one orthography would make things easier. At the same time, some Portuguese politicians see it as a capitulation to Brazilian interests...

What I am interested in is what effet it will have on projects like Wiktionary and Wikipedia. The orthography has been adopted by some and not by others. It is intended that the whole of the Lusophony will adopt this orthography.. So what will it be for the Wiki projects ??
Thanks,
      GerardM

Thursday, May 22, 2008

A better mouse trap

I have had the privilege to participate at a conference on lexicography at the Aarhus School of Business. It was a truly enjoyable conference; I learned a lot, I talked a lot and I even had the pleasure of presenting about Wikis, Wiktionary, OmegaWiki, the need for standards and SignWriting. As each of these subjects are broad enough to present for half a day, it is not strange when people find that one aspect that is of special interest to them, only becomes clear when talking in the coffee break.

There were several presentations that really made me sit up and listen. Presentations about the preservation of knowledge about Tamil crafts, Bavarian dialects, Brazilian learners dictionaries, specialised tools for the teaching of English, electronic pocket dictionaries in Japan .. the list goes on. Really special for me was a talk on Wiktionary.

I have a tender spot for Wiktionary, I know how much it has improved over the years. I know the people, their dedication and the amount of effort that goes into constantly improving what is essentially unstructured data. Given that I am a bureaucrat and admin on several projects, given that I am still running a bot on most Wiktionaries, I have a clue what is going on.

In the presentation the terminology of business terminology was considered. Business, was looked at and the information was found wanting. There are many translations but several were considered to be wrong, archaic... OmegaWiki is a better mouse trap, it does allow for annotations. But in order to be better mouse trap, it has to catch more mice.

The traffic numbers demonstrate the relevance of Wiktionary; it is the biggest resource of its kind on the Internet. This is in my opinion why it is an important resource. When lexicographers want to make an impact, they have to do better. If they want to cooperate, they will find people knowledgeable, interested and enthusiastic. They will have to tread carefully because their scientific reputation is not what will make this work, it will be their ability to listen, to be judicious in their effort and in their ability to share on an equal footing.
Thanks,
GerardM

Monday, May 05, 2008

RobotGMwikt

Several years I have run as a service the interwiki.py function of the pywikipedia bot software on all the Wiktionaries. I had up to six instances of the bot running concurrently on different projects. The principle behind it was simple; when a word in one project is exactly the same, a link is to be created. This function was created by Andre Engels and he helped me out a considerable number of times.

I stopped running the bot when people started to question the algorithm used by the bot. As I run it as a service, I was interested in changing it only when good arguments were provided why it should be changed. These arguments were never forthcoming, the reason why I stopped the running of the bot was because I was told that it is now ran from the tool server.

Yesterday, I learned that the tool server functionality although maybe really clever and efficient is only used to update the English language Wiktionary. Yesterday I was asked by Spacebirdy to run the bot again because there is a need for it. So I asked on IRC if there was a problem and I was told there was not.

Today I was told to change the algorithm because the English Wiktionary wants to create interwiki links to redirect pages and as it is a "community policy" I was told to abide by it. I asked for arguments why this made sense and no good arguments were forthcomming. It was even conceded that it would be best to discuss this in the village pump.

There are good reasons why a redirect should not be referred to by interwiki links:
  • Wiktionary aims to include all words in all languages. A word spelled incorrectly in one language can be correct in another
  • The specific word does not exist yet in the other project
  • The notion of homonymy is ignored
  • There are a substantial number of redirects that exist as a result of a conversion
Interwiki links are one of the few things that connect the different Wiktionary projects. It is essential to consider its use in the scope of the whole of Wiktionary and not within the narrow understanding of single projects. As more Wiktionaries became isolationist in their point of view it became increasingly time consuming and frustrating for me to run the bot. Given that I have in OmegaWiki the perfect solution for the need for interwiki links, I have no need for this functionality in the first place.

What I find disappointing is that the quality of service that I provided is no longer there. It is another fine problem that a community / project council could deal with.
Thanks,
GerardM

Friday, February 08, 2008

When a language is not a language

Kurdish is not a language, Chinese is not a language. They are both macro languages. Macro languages are a way that allow for one ISO-639 language identifier to refer to multiple other ISO-639 language identifiers. Chinese is in many ways easy to understand; these are many languages that are all using the same script, the same writing system; a script that is not sound based. Given that the Library of Congress was the maintainer of the ISO-639-2, it is perfectly understandable that many languages were considered the same, were considered to be Chinese.

Kurdish is distinctly different. Kurdish is also very much a political story. The Kurdish language and the Kurdish people and culture are seen as problematic in all the countries where the Kurds live. Linguistically, Kurdish is considered to be three languages. Central, Northern and Southern Kurdish. These languages are written in the Arabic and Latin script.

In the localisation of MediaWiki at Betawiki, there is Kurdish in Latin and Arab script. The question is, is it Kurmanji or Sorani because I have been told that transliteration is how you get the Arabic version.. I have been told that it is Sorani.

These questions have implications. Take for instance Persian or Farsi, There are successful WMF projects that start with fa. The language is also a macro language; it is divided in Western- and Eastern Farsi. I am quite positive that the fa.wikipedia.org is in Western Farsi. But the url should be pes.wikipedia.org.

There are people that I know and respect that tell me that the difference between Farsi and Dari is only 2 to 3 % and that consequently this difference is an abomination. Other equally respectable people tell me that the difference between Farsi and Tajik is only 2 to 3%. Nobody has so far insisted on bundling the two together.

The request for the fa.wikiversity.org is in the discussion stage. It does fulfill the requirement for localisation. There is a sufficient group of people interested in making this project a success. It is just that I do not feel comfortable with the designation fa for this project that prevents me from changing it to the "eligible" status.

What I want is to have the "fa.wmfprojects.org" renamed.

Thanks,
GerardM

Monday, November 26, 2007

Wiktionary upset

When I look at the Wiktionary website at the moment, it does not show yet that the French language Wiktionary has more articles then the English language Wiktionary. I think it is absolutely wonderful because if anything it shows that you should not take things for granted in Wikis.

Not taking things for granted is a healthy attitude. There is an inherent bias against the French language Wiktionary in the Alexa numbers. However, I am impressed by the numbers quoted.

All the bigger Wiktionary projects have used bots to build up their content. I can imagine that a healthy rivalry will make the numbers go even higher, this would benefit the users of Wiktionary because I trust the Wiktionary communities to watch the quality of the content :)

Thanks,
GerardM

Thursday, November 15, 2007

Statistics

On Alexa, OmegaWiki has for the first time gone through the 100.000 traffic rank. This number gives in Alexa better graphics so it is really welcome news. The daily rank of 99.813 is extraordinarily compared with our three monthly rank of 485.925. It however indicates that something is working in our favour or maybe we are doing something right.

When you compare the OmegaWiki statistics with the Wiktionary stats they are doing really great. With a daily rank of 1.601 it is time for the Wikimedia Foundation to demonstrate that they have a valuable resource in Wiktionary :)

Thanks,
GerardM

Wednesday, July 25, 2007

WiktionaryDev

WiktionaryDev was pointed out to me and there is much to like. I really like the fact that they brought some tables to the party. The system now knows when there is content in a language. I experimented a bit and added both the word Cherokee and a word in that language.

Superb is the possibility to indicate that two labels indicate content in the same language, this brings the content of Greek and Greek (modern) for instance under the same heading. Having indexes for each language is really powerful; it allows people that are interested to work on one specific language.

The WiktionaryDev functionality builds upon the standards that were adopted in the en.Wiktionary. It is absolutely fabulous that the hard work of standardising Wiktionary results in all this new functionality.. :)

Thanks,
GerardM

Tuesday, May 15, 2007

Wiktionary quality issues II

In a previous blog I wrote about the Russian Wiktionary being ostracised by the Polish Wiktionary. There was not enough content and consequently they did not want to have links on the Polish Wiktionary.

Yesterday, I blocked a bot run by an Arab Wiktionarian. I blocked it because it did not comply with the way interwiki links are created, it also did not have a bot flag. A link is only created when the words are exact matches. The problem at the Arab Wiktionary is that they have imported with a bot many words, English words, and they are all upper case.

I have blocked the bot because it is technically wrong. I do run my bot to correct the "damage". The biggest damage however is in the lack of communication between the Wiktionary projects. It is for this reason that I am of the opinion that though courageous efforts, many of them are failures.

Thanks,
GerardM