Showing posts with label interview. Show all posts
Showing posts with label interview. Show all posts

Monday, September 28, 2015

#Wikidata - ten questions about #Kian

The quality and quantity of Wikidata relies heavily on technology. When the people who develop their tools collaborate, the results increase exponentially. Amir has made his mark in the pywikibot environment and now he spreads his wings with Kian.

I have mentioned Kian before and I am happy that Amir was willing to answer some questions now that has over 82,000 edits.
Enjoy,
      GerardM

What is Kian
Kian right now is a tool that can give probability of having certain statement based on categories but the goal is to become a general AI system to serve Wikidata.
Why did you write Kian
Huge number of items without statement always bothered me and I thought I should write something that can analyse articles and take out some data out of articles.
How is Kian different from other bots
It uses AI to extract data, I have never seen something like this in Wikipedia before. The main advantage of using AI is adaptability. I can now run Kian on languages that I have no idea about them. 
Another advantage of using AI is having probability which can be useful in lots of cases such as generating list of mismatches between Wikipedia and Wikidata that shows possible mistakes in Wikidata or Wikipedia.
Is there a point in using Kian iteratively
With each run of Kian Wikidata becomes better. After a while we would have so much certainty in data that we can assure Wikipedia and other third party users using our data is a good thing
What can Kian do other bots cannot do
First is generating possible mistakes and building a quality assurance workflow. 
Another one is adaptability of adding broad range of statements with high accuracy.
What can Kian do for data from Wikipedias in "other" languages
We can build a system to create these articles in languages such as English since using Kian now we have data about those articles.
Let me give you an example: Maybe there is an article in Hindi Wikipedia, we can't read this article but Kian can extract several statements out of that article. Then using resonator or other tools we can write articles in English Wikipedia or other languages.
What question did I fail to ask
Plans about Kian. What I'm doing to make Kian better. Hopefully we would have a suggesting tool using Kian very soon.
What does it take for me to use Kian
We have a instruction in github you only need an account in Wikimedia Labs
Does Kian use other tools
Yes, right now it uses autolist which makes it up-to-date.
What is your favourite tool that is not a bot
Autolist, Wikidata can't go on without this tool.

Wednesday, November 20, 2013

#Wikidata & #Freebase - an #interview with Denny Vrandečić

When Denny signalled the availability of data that links Wikidata to Freebase, there was a lot of follow up and several questions were left unanswered.  So I asked my questions and, I am happy with the answers I received. So enjoy this interview about the Wikidata and Freebase connection. 
Thanks,
       GerardM

By moving to California you embrace yet another culture, language.What does it do to you, your family
My wife is from Uzbekistan, I am from Croatia. We met in Israel. She has lived the last few years in Estonia, I in Germany. We have moved to California, and we are very excited about this move. Finally to live in a country where we both speak the language! And also, we have high hopes with regards to the weather.
The US in general, and San Francisco in particular, is a real melting pot - or rather fruit salad, as they say these days - of people from all over the world. This seems to be a good match to our own background, so we are looking forward to see what this will mean for us.
What is your job description and title at Google (and what is it you actually do)
My job title is "Ontologist". I am working on Google's Knowledge Graph. The Knowledge Graph for Google is basically what Wikidata is for the Wikimedia projects: a repository of structured knowledge about the world, that is used in many different ways in various different applications. I am working, together with a great team, on the schema and ontology of the Knowledge Graph: its data model, schema, and the way it captures and represents knowledge.
What is Freebase and how does it compare with Wikidata
Freebase is the publicly available and editable part of the Knowledge Graph. Freebase and Wikidata are very similar. At the same time, there are quite a few differences: the user communities, the incentive architecture, the license, the way sources and knowledge diversity are handled, to name the big differences. There are also minor differences in the data model, the way classes, properties, types interact with each other, the UI, the workflows, the prominence of internationalization, currently also the size and scope of the knowledge bases, etc. Wikidata currently has 22 Million statements, Freebase has 2 Billion facts. Now one should always be careful with such simple metrics, and especially these numbers are not comparable, but it hints at some differences in the knowledge base.
Do you work on FB or is this only an outward facing project
Freebase is a part of the Knowledge Graph, and I work on the Knowledge Graph, so yeah, I also work on Freebase.
With a link between FB and WD it will be possible data can flow from WD because of its license. How is it for a flow from the other direction
I am not a lawyer, but unfortunately it seems that the licenses used by Freebase and Wikidata would not allow for that direction. We are aware that this is not good, and we are working on ways to fix this.
But even without actually letting data flow from Freebase to Wikidata, such an alignment can be valuable for the Wikidata editors: bots could compare the data, and flag inconsistencies. Bots could do simple comparisons, like, go through all the countries, compare the capitals, and report on the differences. And this can be done not only with Freebase, but with any other structured or semi-structured base that Wikidata connects to. One example is the work that Maximilian Klein did, comparing the sex/gender of authors connected through the VIAF ID.
The data for the FB<->WD connection is based on a WD dump, How will it be maintained in the future 
This is not decided yet. In principle, we could create this dump regularly, but I am unsure if this will be needed. We will have to see how things develop before deciding on what to do next.
When info exist in both WD and FB and it differs, what would you like to see done
The difference to be fixed, obviously! Both knowledge bases may and do contain errors, and the ability to compare them can lead to an increased quality level in both. It is obvious that the respective communities should take care of such errors in their knowledge bases. By providing links between the knowledge bases, this should get easier to automate.
Would that be a model for collaboration with other sources like VIAF
Yes, in many cases that should work. It is clear that some sources and their communities are more open to corrections than others. Wikidata can become a central hub for identity on the Web, collecting and reconciling IDs from many different sources. What makes Wikidata so interesting in itself is its commitment to let everyone share in the sum of all knowledge - both in participating in creating it as well as in accessing it. The barriers are so much lower than almost anywhere else. In which other knowledge base can you easily fix an error?
Does FB have sources for its statements
Yes, but this is quite a different notion from what Wikidata does with sources. When a dataset is being uploaded to Freebase, the source for this upload is usually recorded. It is closer to what is usually called "provenance", whereas Wikidata's notion of a source is closer to what is called a "reference", an external authority that makes the given claim. A strength of Wikidata is the diversity of the sources, and that they can be refined later by the community. A statement in Wikidata can have several sources (which makes sense if you think of them as references supporting the statement), but in Freebase every statement only has one source (which again makes sense if you think about it being the provenance of that statement, from where this statement came from).
spot the differences
What do you prefer, no diffs or sourced diffs
(I don't understand the question, and rephrase it to: do you prefer an unsourced statement over no statement?)
In most cases, yes, I'd rather have an unsourced statement than no statement at all. In a perfect world, all statements in Wikidata would have great, authoritative sources. But for now, I really think that Wikidata should be lenient with regards to unsourced statements in most cases. There are obvious cases where this is not true: data about living persons has to be more carefully sourced, especially when it has the opportunity to hurt the person. Also, we have to remember, that a Wikipedia community can decide to use a statement from Wikidata only if it is sourced, and to drop it otherwise. Once Wikidata has matured a bit, I expect it to move towards a more stricter policy, and to see tools develop that help with getting there. That will be a quite exciting time.
What FB feature would you LOVE to have in Wikidata
The expressive query features of Freebase. The Wikidata team is hard on working towards enabling this, and I am very much looking forward to see it happen: this will open a whole new world of possibilities for everyone using Wikidata's data, and it will lead to much more visibility of the data and help to refine the Wikidata properties.

Wednesday, October 30, 2013

Magnus, master of #Wikimedia #tools

When Magnus agreed to answer some questions, I wanted his answers to have impact. So I asked a few influential people what they would ask Magnus. I am happy both with the questions and the answers; they offer us plenty opportunities for making more of an impact.
Thanks,
      GerardM

Lydia: what's his next cool hack
Technically, since the moment I received these questions, it would be AutoLists. 
I haven't decided what to do after this. There is the prospect of extending/rewriting the Free Image Search Tool to use Wikidata, and to suggest images for Wikidata from the various Wikimedia projects in return. 
Also, with the long-term plans to use Wikidata, or at least its software, to finally get a grip on metadata for files on Commons, it might be interesting to create and seed a database for the "obvious cases", which could then in turn seed "Wikidata Commons", when it finally shows up.
Lydia: and what he'd like to do but doesn't get to (so others can do it)
A million things! Many of my tools could do with some fixing, code review, extension, you name it. If you can code, have a look at the existing tools and maybe sign up for a tool or two which you would like to improve (not just mine, of course!). 
I still have to port tools from the toolserver. Have a look here at tools to port and, help me prioritise! 
Another thing I had started back on the toolserver was a "tool pipeline", where you could chain up tools so that the output of one tool is the input of the next. This could be quite powerful, especially for less tech-savy users who can't write their own glue code, even if the concept sounds uncomfortably like the human centipede ;-) Have a look at my attempt back in 2012.
Brion: What's the awesomest tool you've created that people don't (yet?) know about?
That's a hard one, as I carefully and thoroughly spam the community for every new tool I create ;-) 
One that is probably less known, and that I use myself on Wikipedia, is a special CSS for wide screens, where it shows the Wikipedia main text in a central column, and arranges TOC, infoboxes, thumbnails, and small tables on the left or right. Screenshot & CSS

As for actual running code, less known but with huge potential IMHO, are live category intersects on Wikipedia. Together with [[User:Obiwankenobi]], we came up With this: Screenshot JS
Brion: What's the best way to make things happen? (getting cool projects done, even if nobody else is working on them yet
Have something to show people. Even if it's little more than a mock-up. Even if it's full of bugs, slow, ugly, fails half the time, and does only 1% of what it should do. A single thing that people can see and click on can get people to the cause, where a hundred pages of text will not. Stallman's GNUpedia existed on a mailing list with many a fancy idea; Wikipedia was there to see, to "touch", to edit. Even if it was a single, bare Perl script.
Alolita: what are your recommendations are for the next 5 innovative tools WMF needs to fund to build for Wikimedia websites
Not sure about funding for tools; coding projects big enough to fund should probably be designed as Wiki(m|p)edia-default extensions, or even MediaWiki core. In terms of new functionality prioritised in-house beyond the current scope, these would be the most interesting ones for me:
  • OSM integration. This has been on the WMF back burner for years.
  • Commons file metadata as Wikidata. On Wikidata itself, or added to Commons itself.
  • More data for Tools Labs. View data so that we can support GLAM projects; there has been some talk about this lately, but it has been a few weeks away for years now. Also, faster access to the wikitext than through the API/dumps would be useful.
  • Add Wikispecies to Wikidata. It might revive Wikispecies (which seems sadly underused), or at least rescue much of its information onto Wikidata.
  • One old pet idea of mine: Wikimedia (bi-)weekly magazine. Consider an amalgamate of Wikinews (which, let's face it, has pretty much failed as a daily source of news), Wikipedia "did you know"-type interesting articles, interesting places from Wikivoyage (which appears to be booming), fascinating images from Commons, an excerpt from Wikibooks, a short story from Wikisource, you get the idea. An editorial text, or political comment (relevant to the Wikimedia world; think CISPA), from individuals might not be too far fetched an idea. An issue would be generated in the usual fashion, maybe with an editor-in-chief (rotating? elected every three months?) to address the time-critical nature of the project. A completed issue would then be generated in several formats; on-wiki, PDF, etc. Maybe even go through proprietary distribution channels like Amazon or Apple, if the license situation allows it. I am sure tablet users would flock to it, and we might get some of these "consumers" to participate in one of the Wikimedia projects, maybe with "photos by readers" as an entry-level drug :-)
Alolita: How WMF can help support the development community better * process recommendations * policy recommendations * participation in helping grow more developer contributors
I'll answer these as one question. For tool developers, Labs is already a fine solution. There are a few rough edges which could be smoothed, though; "find someone on IRC" is often the way things are done, which is not always the best. There were a few outages recently, which can happen, and it took hours for someone to have a look a the issue, which is not exactly a life-or-death situation, but still frustrating for tool authors, and not just because of "your tool is broken" messages piling up around them. I understand that only core systems can have someone on standby 24/7, but maybe Labs issues could bubble up the chain automatically if no one from the usual Labs crew is around. 
Getting more people to make use of the great Labs facilities is important. Part of that could be a simple visibility issue: if more people knew that Labs offers not only CPU and storage, but live mirrors of the Wikimedia databases, or that they could easily help fixing or improving that vital tool, we would probably see more widespread adoption of Labs by programmers. Labs-based git repositories for each tool could also enhance visibility, and improve code reuse (which I practice heavily within my own tools). 
In terms of adoption of tools and extensions by volunteer developers into Wikimedia wikis, MediaWiki core itself, or as stand-alone systems, clear and time-limited review procedures would be required, as well as assistance by Wikimedia developers to get code up to MediaWiki standards. On the other end, a public listing with project ideas that, if implemented, would be considered for adoption by Wikimedia, could focus volunteer programmer participation. I am talking about a simple list of "we would like to have these, but don't have the manpower" issues, written by Wikimedia staff, as opposed to obscure feature requests in bugzilla that might be long obsolete.
Open source is also to work together on software, tools. Your tools are available, how do we get more people to cooperate on them
By refusing to fix bugs, we could force the bug submitter to do it ;-) 
On a more serious note, I think improving visibility, as mentioned above, of the great things Labs offers, and that participation on existing tools is encouraged and welcome. Just like many people use Wikipedia but still are not aware that they could edit it themselves, potential volunteer programmers might think that these tools are done by Wikimedia staff, or some carefully chosen group, rather than by run-of-the-mill geeks like them.
When you had someone to work on improving ONE tool; what would you want him or her to do?
Extend Wiri a.k.a. "the talk page". It is not so much a normal, "useful" tool, but a nice technology demo, not just for Wikidata, but for a confluence of technologies (data from Wikidata, speech generation from third-party open source, speech recognition from Google) enables functionality that plays in the same ballpark as big commercial solutions like Siri or Google Now. 
I am well aware it will never be en par with the Mathematica/Wolfram Alpha backend in terms of calculation power, or up-to-the-minute sports results. But even simple improvements, like a listing of options if the (speech) input is ambiguous (e.g. several people with the same name), translating complex questions into Wikidata query requests, displaying images, maps, or using AutoLists to show multiple results, could turn it from a mere demo into an actually useful tool. With proper, fault-tolerant grammar parsing, real reasoning, and the ability to explain the reasoning, it could be amazing!

Wednesday, October 23, 2013

''#Pywikipediabot'' ''20:00 Thursday, October 24, 2013 - Sunday, October 27 22:00'' UTC

The one tool used by most bots has not only stood the test of time, it is getting ready to complete its rejuvenation. It moved closer to the tools used in the Wikimedia Foundation. The road map has it that a triage is needed.

This interview happened largely on the pywikipedia mailinglist. Another great resources for bot runners. What you read is a compilation, more takes can be found in the list archive.. :)
Enjoy,
      GerardM

What is pywikipedia?
  • A Python-based framework to manipulate Mediawiki installations. Any installation, not only those run by the Wikimedia Foundation, can be worked on. You can change everything, you can change in the Wiki as on  an editor also per API. Thus, pywikipediabot can be used to create so called bot tasks, to do the same change to a lot of pages and do it fast.
  • A huge bunch of scripts which use the framework above for all tasks you can think of: One script for example for mass uploading pictures (like scans of a book), one script for cleaning up page source code (like removing <br>, multiple empty lines, reorder  interwikilinks,...), scripts for fixing common errors, and so on. change everything you can change in the Wiki as an editor also per API. Thus, pywikibot can be used to create so called bot tasks, to do the same change to a lot of pages and do it fast.
As I understand it there are two versions, "core" and "compat". What is pywikipedia core and what is compat?
Compat is old. Core is the redesign with complete new and cleaner data structures. Most new API functions (like those to modify Wikidata) are only or much better supported by core.
Why keep them both, it must be a lot of work to have to maintain them both..
There are so many working scripts which do their job now for years - thus, there is not much pressure to move them to core. Nonetheless, if you plan to do something new, use core, it's much more appealing.
Recently all the bugs have been moved to bugzilla ... What is it that you hope to achieve by this?
Handling bugs will be a lot more easier, there will be more eyes to keep on debugging process, tasks will be centralized, it will be more wikified by merging in WMF infrastructure, you can see more in this suggestion page.  
At Sourceforge you have your code repository and - unlinked - your bug and feature request system which over the years get filled with a lot of dormant, outdated bugs, which were fixed a long time ago. We hope to get a clean list of bugs.
Recently all the code has been moved to GIT ... What is it that you hope to achieve by this?
GIT allows you to work much more naturally and doesn't break everything when you merge your different ideas again. So many new features compared to SVN that I do not want to outline this here. In short: In SVN you can branch your code if you do something different. But merging does never work, therefore, noone uses this important feature. With GIT, merging works (as it stores much more information as SVN). Thus, whenever you start a new task, you branch, you have a new idea.
What is the biggest challenge running the pywikipedia bot?
On the pywikibot side, learning its limitation and working around them. On the larger, bot side of things, performance. You have to always balance thread-safe generators and better support for sections (such as retrieve the whole page but submit a section or vice versa). The API/server performance with the local computer's performance (mostly disk access for me, might be RAM limitations when running on a VPS) and the network performance. While not related to pwb, this is something you have to consider for each new bot one writes. pwb could help by providing more advanced support for threads, such as retrieve the whole page but submit a section or vice versa)
How many people are using the pywikipedia bot and how many people are developing code for pywikipedia bot
It has been widely used since 2003, and It has more than 100 authors but now they are just five active developers. After the launch of Wikidata many bots lost its work :-) e.g. one bot was active on all wikipedias, now it works only on 2-4 and the wiktionaries.
It is possible to use the pywikipedia bot in so many ways... Is it easy to learn what it can and cannot do?
Yes: Just join the irc channel and ask. For nearly all tasks, some "swiss army knife" exists, and the developers are extremely helpful. Just ask.  
When script have documentation inside, its easier. Some scripts have documentation on mediawiki.org, but nobody really knows all about it.
How long does it take before pywikipedia bot supports a new Wikidata data type ?
Usually it takes a week or two based on how much the developers are willing to do. The following datatypes are already supported: item, string, URL, coordinates. Time is nearly done. - > ~ 2 month
Tomorrow there will be the pywikibot triage... What is it and who can participate?
The bug day will focuse mainly on categorizing and prioritizing and closing non-reproducible bugs and fixing them if the bug is not a big deal. Because we just migrated there are ~700 open bugs (some of them are really old and it's fixed long time ago) so we need to clean them up. It's really necessary for us to have this "bug triage".
The good news is that anybody can participate.

Friday, October 18, 2013

#SignWriting, #sign languages - an #Interview with Valerie II

Half of the people who sign are not deaf. Therefore the number of people who benefit from a sign language that can be written is more than just the number of people who are deaf and sign.

Thanks to SignWriting, sign languages can move on from only having an "oral tradition". I have been privileged to witness as the SignWriting community moves slowly but surely ahead in gaining recognition by making their languages and their cultures equal to any other language in the age of Internet.

Knowing Valerie is an inspiration. She is a real mover and shaker for so many people. I rate her as highly as Jimmy Wales. Anyway, these are the questions I put to her. Enjoy!
Thanks,
       GerardM

How do you explain what it means when a language cannot be written?
If I understand it correctly, most of the world's languages do not have a written form. All languages CAN be written. But most languages are not written.

Sign languages are now written languages!  ;-))

But it takes effort by people who know their languages to want to develop a way to write it and lots of languages, for example, in Africa and Asia, may not have the political position, nor the funds, to invest in the development.
How many languages are written in the SignWriting Script?
That is also hard to say, but we estimate that small groups of people are writing their sign languages in around 40 countries…based on real written literature and also by word of mouth…we have definite proof for many, and some proof for some…. Some sign languages have 100s if not 1000s of written documents - ASL and DGS are two examples
Can you recognise what sign language it is from a written text?
Yes, between written ASL and written DGS (German Sign Language) we can definitely see a difference immediately - Other sign languages, not so much yet, because we do not have that much experience comparing written documents - as each sign language has more and more literature, it will be easier to recognize differences in the sign language literature - a lot has to do with the choice of style of writing… German Sign Language is written with more mouth movements than our ASL written documents… so one can see which language it is quickly
What are the most active sign languages (that are writing with SignWriting)?
American Sign Language, Argentine Sign Language, Brazilian Sign Language, Catalan Sign Language, Czech Sign Language, French-Canadian Sign Language, French-Belgian Sign Language, French-Swiss Sign Language, French Sign Language, Flemish Sign Language, German Sign Language, German Swiss Sign Language, Italian Sign Language, Jordanian Sign Language, Korean Sign Language, Maltese Sign Language, Nicaraguan Sign Language, Norwegian Sign Language, Polish Sign Language, Portuguese Sign Language, Saudi Arabian Sign Language, Slovenian Sign Language, Spanish Sign Language, Tunisian Sign Language and more…

The use of SignWriting is growing rapidly. How do you know about how it develops?

Through the internet, in different ways. And through individuals sending me documents. And through mentions on the SignWriting List. Some write documents publicly in SignPuddle. Some get their school degrees - Ph.Ds and Master Degree theses are posted written on SignWriting or using SignWriting… Papers are presented about SignWriting and they are listed in publications, and occasionally people write to me privately or join the SignWriting List or Facebook or Twitter… but I actually do not know how many people use SignWriting and I never will, because it is free on the Internet and the way it spreads is like it has a life of its own
Can SignWriting be used on mobile phones or is there an app for that…
Yes, we have two apps for the iPhone…one from Germany and one from California:
SignWriting App from California, by Jake Chasan (age 16 ;-)
Signbook App from Germany, by Lasse Schneider, University of Hamburg
Are there many schools where they teach SignWriting
SignWriting is spread freely on the internet, so lots of people learn SignWriting on their computers, but there are schools with official courses too, such as

  • Osnabrück School for the Deaf in Germany
  • A school in French-Canada (Quebec)
  • Schools in French-Belgium
  • A school in Flemish-Belgium (Brussells)
  • A school in Poland
  • A University in the Czech Republic
  • A Catholic School for the Deaf in Slovenia
  • Bible Translators teach SignWriting in Madrid
  • A university in Barcelona (Catalan Sign Language)
  • ASL classes in a hearing high school in Tucson, Arizona
  • ASL classes at San Diego Mesa Community College in San Diego, California
  • ASL classes at UCSD, University of California San Diego
  • Courses on SignWriting are taught throughout Brazil, by Libras Escritas, also online
  • A School for Deaf Children in Brazil: Teacher Sonia Messerschimidt
  • Santa Maria-Rio Grande do sul - Brasil , Escola Estadual de Educação Especial
  • Letras LIBRAS , UFSC Florianópolis, (SignWriting is an integral part of the curriculum)
  • the list goes on and on ... 
How hard would it be to have the pupils at these schools write two articles a month ... How many Wikipedias could be started that way…
Just as soon as Steve Slevinski and Yair Rand are ready for us to move Nancy Romero's 37 articles, and Adam's 2 articles and Charles Butler's 1 article, from the Wikimedia Labs to the Incubator, and we have tested the ASL User Interface, and we have tested the new features like linking and selecting text, and when Steve has completed the new SignWriting Editor program that will make it possible to write articles directly on the Incubator site, then of course we can ask teachers and students to test our new software and start writing articles… our software isn't ready yet though. 
That is why Nancy Romero, our most prolific English-to-ASL translator and SignWriter, is writing enough articles to lay a foundation so we can get started - we could ask students to write the articles over in SignPuddle Online, and then we can move those articles over to the Incubator for them - but until the Editor software is completed it wouldn't be as much fun for the students as it will be later - that day is coming and it is an excellent idea for the future.
Why is Wikipedia strategically important for getting more people to know about SignWriting.
I consider it VERY important because it provides us literature for readers to read written in ASL that are not children's stories or religious literature. Ironically we have plenty of children's stories written in ASL, including Cat in the Hat, Goldilocks, Snow White and others…and Nancy Romero has written close to the entire New Testament in ASL based on the New Living Translation, and another Deaf church has also written much of the Bible - but we are really grateful to have articles to read in Wikipedia that are general non-fiction - educational, historical, scientific.
Wikipedias are also important because they will encourage others to write articles in ASL which will indirectly teach people how to write ASL. Another wonderful indirect result will be an added respect from the general public. Surprised visitors will realize that ASL and other sign languages can now be written -
Val ;-)

Thursday, October 17, 2013

#SignWriting, #sign languages - an #Interview with Valerie I

The biggest challenge for a #language to gain permanence is to be written. Many languages became written languages by adopting an existing script. Sign languages are fundamentally different and therefore there was no script to adopt.
Valerie Sutton, a ballerina, developed a method to register dance movements. Linguists who researched sign languages asked Valerie if this could be applied to sign languages. Many iterations later, SignWriting became an ISO recognised script, it is known to be used by at least 40 sign languages that all may gain their Wikipedia.

When I asked Valerie to answer ten questions, her response was to ask the SignWriting community for their opinion. This, the first part, explains the need for sign languages. As it is the most often asked question about sign languages it deserves a full response.
Enjoy,
      GerardM

Some people say "Why do they not use English?" .. How different is a sign language from a spoken language?
A lot of Deaf people DO use English … ;-)

Deaf people who use a sign language as their primary daily language, also use English as their second language. But, they cannot hear their second language, English, and American Sign Language (ASL) and other sign languages are rich languages that give a deep communication that is more profound than speaking a second language you struggle to hear, or cannot hear at all… Lip reading does not give all the sounds made on the lips, and many conversations have to be guessed at most of the time… They say that at best, lip reading gives 30% understanding and everything else is guessed at….

So if a signing Deaf person, whose primary language is American Sign Language, lives is in the United States or English-speaking Canada, they have to get around, and they learn English to get by…

Just as I learned Danish when I lived in Denmark. Learning a second language is a requirement and you do your best… But there was one difference for me - I can HEAR Danish, which was my second language years ago when I lived in Denmark… I would have found it much harder to learn Danish if I were Deaf and could not hear it…And truth be told, if I really wanted a deep and profound communication, I would always migrate back to my native language English. My second language did not give me the true communication of my native tongue.

So the question is really not "why one language is better or easier that the other?" but instead a realization that deafness creates a barrier to learning spoken languages, and that first and second languages are different experiences… Deaf people are not choosing one language over the other, but instead managing the best they can between the hearing and Deaf communities.

Both signed and spoken languages are good languages and should be equally respected. And it is my feeling that everyone should learn another language if they can. Hearing people who speak English as their native language oftentimes enjoy learning American Sign Language, but no one asks them why they just don't use English?! (at least I hope not ;-)

In school we are asked to learn a foreign language…so why not learn American Sign Language or other signed languages?

And in return, Deaf people spend most of their lives learning spoken languages to the best of their ability and I give them my utmost respect for the hard work I know they must go through everyday...

Some Deaf people are born into Deaf families. Deaf children in Deaf signing families have a native language, sign language, from the beginning, so their language development is early and considered the same as a hearing child's language development. Spoken languages are a "second language" to everyone in the Deaf family. So they do not feel "different" than their parents or siblings.

Deaf children born into hearing families sometimes have a harder time, because oftentimes the family doesn't even know the child is deaf until later, and so language development may not start early, and also they are different than their own parents and family members.

So native signing Deaf people have their own native language, a sign language, and yes, the grammar and structure of American Sign Language, for example, is quite different than the grammar and structure of English. Verbs are conjugated differently, adverbs and adjectives are in different positions in the sentence, and there are elements of American Sign Language that are much more sophisticated than in English, or at least expressed very differently, and so oftentimes there is not a real way to translate between the two languages that is a true "match"… what takes a paragraph in English can be expressed with a short phrase in ASL, and vice versa… Some say that the grammar of ASL is closer to Russian or Spanish than it is to English...

That is why SignWriting is important. When both languages can be written, both languages can be compared, and understood better…

Val ;-)

Monday, October 07, 2013

More heady stuff about #Wikidata and ontologies

When I asked Emw many questions, I received only three answers. Having an opinion on Wikidata and expressing it is hard. It feels very much like exploring new frontiers. The questions did not go away so I asked them again and, I am mighty pleased that Antoine Isaac was willing to provide me with some answers. Antoine does not do Wikidata, he is a/the scientific coordinator at Europeana.eu. He wrote an email to the Wikidata mailing list that got me interested in asking him my questions. I hope we will collaborate for our GLAM data well with Antoine and Europeana.

I exchanged several emails with Antoine but the answers do stand on their own. I will react in a follow up blog post. I hope you appreciate what Antoine has to say as much as I do.
Thanks,
     GerardM

The Wikipedia article calls upper level ontologies "political" why should Wikidata be interested in any of that?
I wouldn't call them 'political', that's a bit far-stretched. Indeed ULOs embody some abstract considerations on how to represent the world. And as said before, this can backfire as soon as you consider an open web of data, where different representations may well co-exist. Considering a single upper-level ontology as the guiding principle for everything is dangerous. But it is true that they provide valuable bodies of knowledge to re-use, and it may make sense to re-use them for specific domains (e.g. biology, geography) that could be compatible as a whole with the approach of one ULO.
Does it not make sense to group statements together as qualifiers as part of a statement (like it is done for office held [1]) ?
I like qualifiers and dislike them at the same time!On the one end, it's good to have some meta-metadata about the provenance of a statement (who made/endorsed it), or its scope (e.g. the time it applies). In your example, this is the "start date" and "end date". This practices actually fits what is happening in the RDF / Semantic Web area, where a lot of work is being done about Provenance, and many people use quad stores with 'named graphs' instead of just triple stores. 
On the other end, I am anxious about qualifiers being used with other semantics that "here's some info about a statement". In your example, "preceded by" and "succeeded by" are more difficult to interpret in this sense:  (compare with property P580 that mentions explicitly "statement"). I mean, it is possible to interpret your qualifiers as data on statement. But I really feel that people (and you?) will understand it as, say "Te Rata is the predecessor of Korokī Mahuta" and not "The statement 'Te Rata held the office of Maori Monarch' is the predecessor of the statement 'Korokī Mahuta held the office of Maori Monarch'". Which should be the right thing to do (I mean, the one compatible with the "start date" semantics).  
Note that what you told in the other email is the kind of use of qualifiers that would worry me: it is the person who has a birth/death date or a sex, not the statement 'is a person'.Of course one could see the birth and death date to influence of the 'date of validity of the statement "is a person"'. But still we'd be talking about two different things, from a knowledge representation perspective. 
Note also (just for the fun of refering to upper-level ontologies and their dangers) that for some ULOs 'is a person' doesn't have a begin and end date. Being a person is 'rigid' i.e. it must stick to the subject forever. You can't have been a person once and then cease to be a person. Even if you die you're still a person...And don't tell me that rigidity foresees the some ontologies may have Person an an anti-rigid property. This may probably not be the choice made in your favourite ULO. Unless it's one that addresses both reality and beliefs as two possible sides of a same property. But then, good luck re-using it!!!
DBpedia does not have qualifiers, will this impact their ability to use data from Wikidata?
I can't really speak for DBpedia. But I'd say that if qualifiers are used in a way that is both consistent and compatible with the understanding of 'named graphs' in RDF, then they might be interested.
As we map Wikidata items to the content in other repositories, what do we need to compare the data from these repositories
Which repositories are you talking about? Which content? Are 'repositories' knowledge bases, like OCLC or Europeana in the book/artwork domain? Is 'content' 'data'? If yes, then what is required to establish correspondences is hard work looking at the fields and seeing how they correspond. This may be of course alleviated if all (wikidata included) look at what's already happening and try to minimize the risk of coming with data models that are too indiosyncratic. (this is why something like RDF is quite useful!)
When differences in content between repositories are found is there a standard method to harmonise the content
Again assuming a reading of 'repositories' and 'content' as above. I don't think there is a standard method. What you can prey for (and of course as designers of data repositories, we are somehow in the position of making it happen!) is that all repositories keep track of as many unambiguous identifiers they can keep track of (e.g. ISBNs for books) which would help automatic reconciliation. Otherwise make sure the data on the content (where 'content'='the object in the real world') is as complete as possible. For works of art that would mean that comparisons can be made on titles, creators, dates and place of creation, etc.

Saturday, October 05, 2013

10 #Wikidata questions for Lydia

Lydia is the new project leader for Wikidata. She has already done a great job communicating for the Wikidata project and, with much of Wikidata well established, communication will increasingly be the critical success factor. The Wikidata team works together well so I think she will do really well. Enough reasons to ask her some questions. I hope you will enjoy them as much as her answers.
Thanks,
     GerardM

Wikidata exists, it is being used. Do we know how much it is used?
Unfortunately not. Of course we get to see some great uses but not the whole picture. I love seeing all the different ways the Wikipedias are already now using Wikidata that I know about. For example the Occitan Wikipedia building whole infoboxes based on data from Wikidata. Or the English Wikipedia that compares its local IMDB identifiers against those in Wikidata and puts pages into a maintenance category if they do not match so a human can look into it and fix it. But this is only inside the big Wikimedia projects. At the same time people are building 3rd-party tools that use data from Wikidata. The most recent and very impressive one is the Wikidata tempo-spatial display. We are seeing more and more of these pop up and I am looking forward to what other useful, cool and even crazy things people will come up with.
You want people to trust Wikidata data. How do you envision this to work?
Wikidata has just started and we're seeing people add a lot of data to it for all the world to use. This is fantastic. At the same time we need to make sure that the data in Wikidata is reasonably trustworthy. I say reasonably trustworthy because in the end this is just like Wikipedia. It's not 100% perfect but we're all making a huge effort to keep it as accurate and correct as possible and people have come to rely on this. Wikidata is in a bit of a better situation here though than Wikipedia for a few reasons. First of all since the data is structured and machine-readable it is much easier for a computer to find inconsistencies and alert an editor about them. One example would be that a political office is said to be held by X but X is said to be an animal. Now there were probably a few cases where a political office was held by an animal but in general things like this should be flagged for an editor to reexamine it. And there are many such cases that could be checked. Here you find a few such checks that are already in place. The other advantage Wikidata has is that it will be much easier to verify a given data point against an external source once that is given in Wikidata. And the third advantage Wikidata has is that it will be watched by potentially a lot more people. Changes in Wikidata show up on the watchlists of all editors who are watching the corresponding page on their Wikipedia as well as the recent changes of that Wikipedia. So all in all we are in a pretty good position. However we are not there yet. More tools will need to be developed, existing ones improved and most importantly people will need to spend time adding sources to statements in Wikidata.
When you talk about the user experience of Wikidata, what are the limits of this user experience as far as you are involved?
I want Wikidata to be a joy to use and I want it to be easy to use. At the same time experienced editors need to be able to navigate the site and complete their tasks quickly and efficiently. We will have to find the right balance there over and over again. The same thing goes for all the missing features we still have to develop or roll out - queries and the numbers datatype for example. I will put more emphasis on user experience but moving the project forward on a feature level is also very important.
One obvious target for Wikidata is to include all the information contained in info boxes. How far off are we before Wikidata can service most subjects that have info boxes.
I think the big missing piece of the puzzle is the numbers datatype. We're not too far from rolling out a first version of it. Once we are able to also deal with units we are well on our way to that target.
You want people to better understand Wikidata. Is that not even more complicated than understanding templates and info-boxes?
They should absolutely not need to understand everything about Wikidata. This would not work. But for those who interact with Wikidata regularly it should be easy to understand what is going on - not in detail but the bigger picture. To get there we need to improve a few things in the user interface but we will also need to adapt our help pages to be less technical.
What can people do to make Wikidata be useful in their language
Wikidata is inherently a multilingual project. It allows you to use the site in your language and will show you the data it has in your language. Have a look at for example item Q2 in German, English and Spanish. To be able to do this Wikidata needs to know the names of all these things in the different languages. That's what we call "labels". These labels are really important in all kinds of places in Wikidata as they make it possible to refer to things by their name instead of the identifier we gave it - Q2 in the example above. So the best way to make Wikidata more useful in a language is to enter a lot of labels and descriptions in that language. The Special Pages Entities without Label  and Entities without Description are there to help with this as well as the Terminator tool that sorts them by how often they are used to increase effectiveness. This is especially important for the smaller languages. If you speak several languages you should also add a babel box to your user page on Wikidata like I have done on mine. Once Wikidata knows which languages you are speaking this way it will show you the labels and descriptions for a given thing in these languages as well and will let you complete them in case they are still missing. Over time Wikidata will become more and more useful in all languages we support. One of Wikidata's major goal is supporting smaller Wikipedias. This is one of the most important steps on the way to get us there.
People like Magnus and Denny visualised data that exists in Wikidata, how important do you think visualisation is?


Visualisations are crucial for Wikidata. They are a way for us humans to make sense of the vast amount of data. They allow us to see patterns. They allow us to see where we are missing information. They allow us to see where we have outliers in the data that need closer examination. (Belgium has the information that it shares a border with Australia? Probably not when you look at it on a map...) These things are a lot easier to spot and make sense of when you have a nice visual representation of them. But they also show us how far we have already come and give us a sense of achievement. Look at this gif for example. It shows the progress of adding geocoordinates to Wikidata over the first days this was possible. And last but not less important visualisations are of course beautiful and fun.
What can be done to help people be effective in Wikidata
Make it easy for them to get started and understand the basic concepts quickly. Make it easy for them to find like-minded people for example in the task forces. Keep the number of rules low. Create more tools like Terminator and more Database reports to make it easy to find the areas that need more work.
Do you consider that Wikidata is a project in its own right or is it beholden to the Wikipedias, particularly the biggies?
Wikipedia is definitely the most important use-case for a long time to come. However as we're already seeing now Wikidata's data is of use for many many parties outside Wikimedia as well. This will only increase. It is definitely a project in its own right.
What is your dearest memory of Denny as your predecessor?
I have many dear memories. When we met for the first time a few years ago it was in a small room at our university in Karlsruhe to discuss Semantic MediaWiki and its community and development. It struck me that both Denny and Markus (the two founders of Semantic MediaWiki) just got it. They understood what it means to build a community and develop a project in the open with all its benefits and drawbacks. A very rare trait I can tell you. Since then we've built this amazing project and an incredible community gathered around it. Along the way we've been to Wikimania in Washington, D.C. and Hong Kong and many other events. At each of those events we've met incredible people who are passionate about what we're doing and willing to help. I'll never forget that and I'm looking for more of it to come.

Friday, September 27, 2013

Importing data from the #Polish #Wikipedia

All new information in #Wikidata has an origin. It can come from many sources and the quality varies. When Matmarex mentioned his source as a Polish project about information about persons, I wanted to learn more.

This is the kind of project that we should welcome at Wikidata. Please have a read and be happy with great undertakings like this.
Thanks,
       GerardM

What is the data you are importing based on
The data is based on the index of biographies maintained by hand by a few dedicated Polish wikipedians at Noty_biograficzne . The nature of created-by-hand data is that not all of it can be automatically parsed, but it was surprisingly consistent – accepting just several common variants resulted in over 60 000 items the bot could understand and only several hundred that could not be used (I am hoping to sort these out by hand). Some typos in the source are unavoidable, but overall the quality seems to be very high.
Getting quality personal information has been a project on the pl.wp for quite some time, can you explain what it means ?
I am not sure myself what the index was intended to be, not yet being a wikipedian when it was started in 2004 – possibly a crossover between a category system (the concept of a category was only introduced on that year, I don't know what was first) and a list of articles needing creation. Currently it serves as, well, an index – list of all biographies on the Polish Wikipedia, ordered alphabetically by last name (or, in some cases, by pseudonym). It's easier to find what you're looking for if you only remember last name of a person and possibly their occupation than using the built-in search system (for example search suggestions are ordered by article title – thus first name – and it's not possible to limit results to only biographies).
You are running a bot adding descriptions in Polish, what software are you using..
Unlike most bot operators I'm not using the Pywikibot framework – I opted for my own custom-written library in Ruby called Sunflower and a a set of scripts using it.
Does it use the Wikidata API and why is this important
Yes, both for uploading and using the information. The API is basically what made the project possible.
What other data do you have about all these people
The old index contains birth and death years in addition to the descriptions. I didn't upload it because it's basically unsourced (and unsourced data is seemingly as frowned-upon on Wikidata as it is on Wikipedia, if not more) and because, when I tried comparing them with birth and death categories on the biographies themselves, I found over 1500 conflicts.
Can you use your data to compare against the data on Wikidata
Not really; there isn't much to compare in this area, especially since I was uploading basically free-form text. Only several hundred items out of the 60 thousand my bot edited already had a description in Polish, me and a couple other editors reviewed them all in a few days.
The birth/death data could be compared, though, but I haven't looked into it. Any help would be welcome!
Can you add your data where Wikidata has none
There are a few things the index uses that are not yet present on Wikidata. The birth and death dates are the biggest one, real names of people using pseudonyms (such as Sting or Madonna) would be a valuable piece of information as well. I didn't try to upload either – the dates would need better sources (the 1500 conflicts are a strong indicator that information from Wikipedia might not be good enough) and there are currently no properties defined for first / last names because of how complicated the topic is (there is currently a discussion under way).
Did you know that this type of data becomes available on several Wikipedias in stub articles
I considered parsing the articles themselves to extract the descriptions, but decided that this would be too error-prone to automate entirely. Instead I developed a gadget that helps users write short descriptions for biographic articles by extracting the information from the lead-in paragraph and presenting it on the index pages – they can be adjusted by a human and saved to Wikidata with one click! This benefits both projects at once and I think is a good example of how they can work together.
Is it possible to transfer this biography project from the pl.wp to Wikidata
It could be done entirely on Wikidata and using Wikidata information, but there are two preconditions – presence of the required information on Wikidata (birth/death dates, last names for correct sorting) and ability to generate lists from the data (so-called "phase 3" could accomplish this).

Tuesday, September 17, 2013

Some answers about the heady stuff of #Wikidata

I asked Emw questions as a result of his email about the migration away from the "GND main type". I am happy with the answers I received and I hope you will enjoy reading them.
Thanks,
      GerardM

At Wikidata, I contribute to discussions about properties, where I espouse using W3C recommendations and conventions from the wider Semantic Web. I'm also active in discussions about how to model molecular biology data.  Outside of Wikidata, I've been an active contributor to Wikipedia and Commons for several years.

I only had time to answer three of your questions, but I did that much pretty extensively.  The remaining questions are mostly beyond my knowledge and I don't have any well-formed opinion on them.  If you'd like, I can try to answer those questions or others next week.

My answers to your questions:

1) The GND system has been ditched. Can you explain why this is a good thing?
The GND main type property has several major problems. Deprecating that property helps us focus on better solutions for classifying knowledge on Wikidata.
Major issues with P107:
  1. With the GND main type property, "person" can mean things well beyond the common understanding of that word.  It can mean things like Coco Chanel -- i.e. 'person' as conventionally understood -- or it can mean a god, literary character, pseudonym, collective pseudonym or spirit.  The standard response to this glaring issue is "'person' is meant to generalize, don't take the term literally". That is not a sufficient solution.  If a classification system for all human knowledge considers Vishnu and Coco Chanel to be both be 'persons', that's a big problem.  Beyond giving users bizarrely unexpected query results, it means properties that should be safe to assume for any given 'person' item simply cannot be.
  2. Any item that is not a person, place, event, organization or work is classified as a "term", which contains virtually no information.  We need to be able to classify things like gravity, carbon, DNA, cancer, clarinet, Twelver Shia Islam, fashion boot, dog and potato as more than simply "terms".  One sixth of the property is kruft.
  3. Not even the GND directly uses GND main types.  The GND Ontology has a hierarchical class system and the Deutsche Nationalbibliothek -- which developed it -- uses the lowest-level, most specific GND class available for a subject.  This indicates that the GND senses the GND main types are not appropriate to use as they are with P107.
  4. The nature of P107 implies that the property is only for the highest level of classification, and that additional properties would be needed for each level in the hierarchy of classification for lower-level types. This would entail lots of unnecessary work to create and update classifications. For example, want to specifically classify Nauru? If property P107 were to persist, then you would need to add something to the effect of "main type: Place" and "subtype: Administrative unit". The problem gets drastically worse for subjects with more levels of classification, like organisms, instruments, molecules, diseases, towns, etc.
The GND system itself -- the GND Ontology -- is not the real problem.  The real problem is that P107 is a "main type" property.  In a project to structure all knowledge -- which Wikidata is -- restricting all items into a small set of types will inevitably lead to many, many classifications that are either A) too broad to be useful or B) simply incorrect.
2.  You sent an email where you asked for attention for what is to be next. Why should there be something next?
Because -- although it is complex -- the world has structure, and classes or types are a useful way to express that structure.  The lopsided debates in the Primary sorting property RFC indicate that so-called "main type" properties (sometimes also called "principal group" or "primary sorting" properties) are a bad idea.  However, that does not mean that the basic notion of grouping things into "types" or "classes" is also a bad idea.
A much better solution for classifying things is to use "type" properties recommended for the Semantic Web by the W3C -- that is, use rdf:type and rdfs:subClassOf. These properties exist in Wikidata as instance of (P31) and subclass of (P279). These properties have been part of W3C recommendations for the Semantic Web for almost a decade. They are fundamental properties used in large controlled vocabularies to structure data into knowledge.  They facilitate classification at an arbitrary granularity.  Together 'instance of' and 'subclass of' can classify all subjects and be used to determine precisely where each subject exists in the hierarchy of knowledge -- or, perhaps -- a collection of hierarchies of knowledge.

Not only do they solve those structural problems of P107 and other "main type" properties, but by being based on W3C recommendations, instance of (P31) and subclass of (P279) also make Wikidata more interoperable with the rest of the Semantic Web.
That said, deciding on properties like P31 and P279 is only the beginning of forming a better way to do classification on Wikidata.  We need a way to map the information in P107 to use P31 and P279.  That's a topic of active discussion on Wikidata.
3.  The GND is a library system then you mention upper ontologies. What is the difference, and how are they practical in the Wikidata context?
The GND (Gemeinsame Normdatei) authority file is used as a library classification system, but it's based on the GND Ontology.  The ontology has a hierarchy of high-level entities and sub-classes.  The P107 property is based on those so called "high-level entities", which were called "main types" in Wikidata as shorthand.  The main GND types are person, place, event, organization, work, term or "undifferentiated person".  These main types are fine as a way to classify items of general interest in a large library, but they're much too small to form a sound basis for a classification system for all human knowledge. 

That's what upper ontologies are for.  An upper ontology is a way to have standard vocabulary about high-level entities in our world.  The idea is to formalize these very general concepts in a way that captures the richness of human language while also being precise enough to be machine-understandable. 

For example, the Suggested Upper Merged Ontology (SUMO) sets a class "entity" as the most general type of thing -- everything is an "entity".  From there, SUMO classifies things in the world as either "physical" or "abstract".  "Physical" things can be "objects" or "processes".  "Abstract" things include so-called "set-classes", "propositions", "quantities" and "attributes".  (More information on SUMO is available in Towards a standard upper ontology.)
There are several other upper ontologies available, like BFO and UMBEL.  I am not an expert in ontologies, and I have not learned enough about each of them to make an informed statement on their advantages and disadvantages.  However, because they seem to offer unifying terminology for different domains of knowledge, upper ontologies strike me as something worth consideration by the Wikidata community.