Dear Katherine, I loved your presentation at the Berkman Klein Center for Internet and Society. It has much to think about and <grin> it is great that you answer the question you want to answer </grin>.
You address questions like "will we let external organisations use our data for their own purposes". My suggestion to you, us all, is why not use our own data for our own purposes.
The Cebuano Wikipedia is seen as problematic on many levels. It is one of the biggest Wikipedias in number of articles and one of the smallest in the size of its community. Like any Wikipedia, its articles are harvested for use in Wikidata and that brings us to several problems but more importantly in the light of your presentation, opportunities.
Problem: the data used the Cebuano articles are based is problematic
Opportunity: import the data in Wikidata first and first do some curation there.
Problem: the data is licensed under a CC-by-sa license and Wikidata is CC-0
Opportunity: collaborate with the copyright holder and ask their permission to include the data in Wikidata
Problem: when text is generated by a bot, the text when saved in an article is fixed
Opportunity: do not save it as an article but generate the text and maybe cache the text
Problem: other organisations use our data to generate information
Opportunity: we generate information in all the 300 languages where Wikipedia does not have an article
Problem: we have information that has no article in any language
Opportunity: we generate the text and maybe cache the text
Problem: Wikimedia officials indicated that issues like the Cebuano Wikipedia are not relevant
No opportunity; opportunities for all our projects are missed
Katherine, we already generate texts using bots, we already cache our data, we do it for English, we do it for Swedish, Cebuano. Why leave it for the companies of our world to generate text where there is already so much? We can do better, do the same and do it for all our languages as well.
Thanks,
GerardM
Showing posts with label Publishing. Show all posts
Showing posts with label Publishing. Show all posts
Sunday, October 22, 2017
Wednesday, June 13, 2012
The eye on the prize
Noodlot, in translation "Fate" is a book by Louis Couperus. Couperus is one of the literary giants of the Dutch literature and, it makes excellent sense to make this book available to a reading public.
To this end, someone took the djvu file from the Gutenberg project and started the proof reading process at the Dutch Wikisource. This process is still ongoing, I did two pages and I regret it.
The regret is because the book has already been proofread at project Gutenberg. As the book is in the public domain, the proof read transliteration is also in the public domain.
What is the point ? Why not do another book ?
The aim of transliteration and possibly the editing of the lay out of the book serves one aim. Making it available to readers. It makes sense to do the proofreading once. There are many other books that are waiting to be digitised for a first time. There are plenty of other sources that are waiting, waiting in a library a museum an archive.
Let us not waste our efforts. Let us do things once and let us do them well. When we are done with the proofreading, the formatting we need to find a public appreciative of the work done. To make it attractive it helps when there is a lot to choose from and when it is available in the format expected by our intended public.
Thanks,
Gerard
To this end, someone took the djvu file from the Gutenberg project and started the proof reading process at the Dutch Wikisource. This process is still ongoing, I did two pages and I regret it.
The regret is because the book has already been proofread at project Gutenberg. As the book is in the public domain, the proof read transliteration is also in the public domain.
What is the point ? Why not do another book ?
The aim of transliteration and possibly the editing of the lay out of the book serves one aim. Making it available to readers. It makes sense to do the proofreading once. There are many other books that are waiting to be digitised for a first time. There are plenty of other sources that are waiting, waiting in a library a museum an archive.
Let us not waste our efforts. Let us do things once and let us do them well. When we are done with the proofreading, the formatting we need to find a public appreciative of the work done. To make it attractive it helps when there is a lot to choose from and when it is available in the format expected by our intended public.
Thanks,
Gerard
Monday, May 28, 2012
#Language, #script, #Unicode, #font and web fonts
Making the Internet globally accessible is more then running cables. It is also about making sure that you can read and write any language. Once all this is in place, people enabled in this way can share in the sum of all knowledge.
There are few people as intimately involved in supporting languages than Michael Everson. He is known for encoding scripts into Unicode, this requires both technical and linguistic expertise and finally he is a publisher of books written in minority languages.
Enjoy,
GerardM
Are all scripts registered yet ... do we know them all in ISO 15924?
No, not at all. The best-known scripts have been given four-letter codes in ISO 15924, but we tend to be conservative for lesser-used scripts, and try to co-ordinate with proposals for encoding them in the Universal Character Set (a.k.a. ISO/IEC 10646 or Unicode).
Several scripts are not yet encoded in Unicode. Many of them are used by living languages.. How do languages cope?
A script (or character) not encoded can't be used in interchange. People can either use the Private Use Area or hack an existing encoding.
What does it do to the cultures involved ?
The lack of an encoded script prevents a language from using its script effectively in any computer environment.
Is it known how many scripts used by living languages are not yet encoded ?
I don't think we have kept a quantitative inventory. And we always discover something new. I know of a number of specialist scripts like SignWriting and Blissymbols which have not been encoded. We are working on some other scripts, like Woleai and Afáka, and a number of West African scripts, but it is very difficult to contact the user communities to get feedback. There is a huge technological divide. (Not for SignWriting or Blissymbols: for those the problem is a lack of funding to do the work.)
Is it known how many scripts used by dead languages are not yet encoded ?
Again, we don't keep a quantitative inventory. The Roadmaps on the Unicode site are as good a checklist as anything.
Several scripts are encoded but there is no freely licensed font for them. Why is this not part of the process of encoding for Unicode ??
The Universal Character Set is a character set. Both ISO/IEC JTC1/SC2/WG2 and the Unicode Technical Committee work to study character and script proposals, give the characters the right properties, and get them encoded. It is not the function of either committee to establish implementations, or to give them away. The work is already voluntary (and expensive).
MediaWiki supports web fonts ... What relevance does this have for you, what opportunities are there for the Wikimedia communities
It is a great opportunity for Wikimedia to exploit some of the generosity of the many people who have donated to the foundation, and to make good use of the skills of people who have expertise in the Universal Character Set and in font design.
What impact will the availability of freely licensed fonts have on the availability of information in those scripts
For instance, right now anyone viewing any Wikipedia in any language may encounter text in Ol Chiki, or in Runic, or in the simple International Phonetic Alphabet, and pages have to apologize to the reader because their computer may not display the material correctly. This is *bad* for the encyclopaedia.
What difference would it make if the Wikimedia Foundation were to become a player in the development of fonts
People using the encyclopaedia would be able to see the information without worrying about seeing ☐☐☐☐☐☐ ☐☐☐☐☐! From a personal point of view, I can say that at various conferences over the past two years, I have spoken with people in the Wikimedia Foundation, and with people from another very large organization, about this matter -- specifically about exploiting my own expertise in the Universal Character Set and in the provision of rare scripts and characters in web fonts -- yet nothing has resulted. I think the message has got through. But so far no one in either organization has decided to take the necessary principled decision that in order to ensure that the information in the Free Encyclopaedia is actually available to people who use it, complete UCS support should be provided in a suite of freely-available and maintained webfonts.
Provenance is the basis for the establishment of facts. Is transcription in the original script essential ?
Why wouldn't it be? That's the source text. Encoding it correctly means that it can be interpreted by the reader if he or she wishes to consult the primary source. Anything else obliges the reader to use someone else's interpretation. Of course expertise is needed, but the closer one can get to the primary source, the better.
Michael, why "Alice's Adventures in Wonderland" ?
I love languages, and it has been a great honour for me to publish Alice for the first time in a number of minority languages which might otherwise never have seen the text. Alice is available in the following languages: Cornish, English, Esperanto (Kearney), Esperanto (Broadribb), French, German, Hawaiian, Irish, Italian, Jèrriais, Latin, Lingua Franca Nova, Low German, Manx, Mennonite Low German, Borain Picard, Scots, Swedish, Ulster Scots and Welsh and several others translations are being prepared.
There are few people as intimately involved in supporting languages than Michael Everson. He is known for encoding scripts into Unicode, this requires both technical and linguistic expertise and finally he is a publisher of books written in minority languages.
Enjoy,
GerardM
![]() |
| Michael at Chogh Zanbil - Cuneiform .. :) |
No, not at all. The best-known scripts have been given four-letter codes in ISO 15924, but we tend to be conservative for lesser-used scripts, and try to co-ordinate with proposals for encoding them in the Universal Character Set (a.k.a. ISO/IEC 10646 or Unicode).
Several scripts are not yet encoded in Unicode. Many of them are used by living languages.. How do languages cope?
A script (or character) not encoded can't be used in interchange. People can either use the Private Use Area or hack an existing encoding.
What does it do to the cultures involved ?
The lack of an encoded script prevents a language from using its script effectively in any computer environment.
Is it known how many scripts used by living languages are not yet encoded ?
I don't think we have kept a quantitative inventory. And we always discover something new. I know of a number of specialist scripts like SignWriting and Blissymbols which have not been encoded. We are working on some other scripts, like Woleai and Afáka, and a number of West African scripts, but it is very difficult to contact the user communities to get feedback. There is a huge technological divide. (Not for SignWriting or Blissymbols: for those the problem is a lack of funding to do the work.)
Is it known how many scripts used by dead languages are not yet encoded ?
Again, we don't keep a quantitative inventory. The Roadmaps on the Unicode site are as good a checklist as anything.
Several scripts are encoded but there is no freely licensed font for them. Why is this not part of the process of encoding for Unicode ??
The Universal Character Set is a character set. Both ISO/IEC JTC1/SC2/WG2 and the Unicode Technical Committee work to study character and script proposals, give the characters the right properties, and get them encoded. It is not the function of either committee to establish implementations, or to give them away. The work is already voluntary (and expensive).
MediaWiki supports web fonts ... What relevance does this have for you, what opportunities are there for the Wikimedia communities
It is a great opportunity for Wikimedia to exploit some of the generosity of the many people who have donated to the foundation, and to make good use of the skills of people who have expertise in the Universal Character Set and in font design.
What impact will the availability of freely licensed fonts have on the availability of information in those scripts
For instance, right now anyone viewing any Wikipedia in any language may encounter text in Ol Chiki, or in Runic, or in the simple International Phonetic Alphabet, and pages have to apologize to the reader because their computer may not display the material correctly. This is *bad* for the encyclopaedia.
What difference would it make if the Wikimedia Foundation were to become a player in the development of fonts
People using the encyclopaedia would be able to see the information without worrying about seeing ☐☐☐☐☐☐ ☐☐☐☐☐! From a personal point of view, I can say that at various conferences over the past two years, I have spoken with people in the Wikimedia Foundation, and with people from another very large organization, about this matter -- specifically about exploiting my own expertise in the Universal Character Set and in the provision of rare scripts and characters in web fonts -- yet nothing has resulted. I think the message has got through. But so far no one in either organization has decided to take the necessary principled decision that in order to ensure that the information in the Free Encyclopaedia is actually available to people who use it, complete UCS support should be provided in a suite of freely-available and maintained webfonts.
Provenance is the basis for the establishment of facts. Is transcription in the original script essential ?
Why wouldn't it be? That's the source text. Encoding it correctly means that it can be interpreted by the reader if he or she wishes to consult the primary source. Anything else obliges the reader to use someone else's interpretation. Of course expertise is needed, but the closer one can get to the primary source, the better.
Michael, why "Alice's Adventures in Wonderland" ?
I love languages, and it has been a great honour for me to publish Alice for the first time in a number of minority languages which might otherwise never have seen the text. Alice is available in the following languages: Cornish, English, Esperanto (Kearney), Esperanto (Broadribb), French, German, Hawaiian, Irish, Italian, Jèrriais, Latin, Lingua Franca Nova, Low German, Manx, Mennonite Low German, Borain Picard, Scots, Swedish, Ulster Scots and Welsh and several others translations are being prepared.
Thursday, March 24, 2011
#Publishing is what you do to get a public
When you google for a definition of "publishing" the first two results are fundamentally different for a writer seeking an audience.
When something is to be published, many skills are needed to prepare for publication. The material may need editing, proof reading, peer review, presentation, marketing before it is readied for consumption. The skills involved maximise impact and distribution. Choosing a medium for a publication is one of them.
When publishing is not the printing business, maximising a paying public is what an author looks for. Each medium has its own cost structure and each medium has its own public. Getting this mix right and optimising for a return on investment makes the publishing business a business with a future.
Thanks,
GerardM
- publication: the business of issuing printed matter for sale or distribution
wordnetweb.princeton.edu/perl/webwn - Publishing is the process of production and dissemination of literature or information - the activity of making information available for public view. ...
en.wikipedia.org/wiki/Publishing
When something is to be published, many skills are needed to prepare for publication. The material may need editing, proof reading, peer review, presentation, marketing before it is readied for consumption. The skills involved maximise impact and distribution. Choosing a medium for a publication is one of them.
When publishing is not the printing business, maximising a paying public is what an author looks for. Each medium has its own cost structure and each medium has its own public. Getting this mix right and optimising for a return on investment makes the publishing business a business with a future.
Thanks,
GerardM
Subscribe to:
Posts (Atom)



