Showing posts with label GSOC. Show all posts
Showing posts with label GSOC. Show all posts

Wednesday, March 19, 2014

#GSOC - #DBpedia and #Wikidata


The most interesting Google Summer of Code proposal I have seen for 2014 is this one.
4.7. Clean DBpedia datasets and import in Wikidata
The student who will take this task will be responsible for two things:
Clean up the DBpedia errors based on the output of Databugger (http://databugger.aksw.org) or similar. With this the student will generate a more sparse but cleaner dump of DBpedia that will be of general use.
Communicate with the Wikidata community in order to coordinate the import of (parts of) the cleaned datasets and re-use the connections of DBpedia to fetch additional data for Wikidata import.
Mentors: Dimitris Kontokostas, Magnus Knuth (co-mentor)
As far as I am aware they are still looking for a candidate to choose this project. This project is totally relevant and it is exactly the kind of repeatable process that Wikidata should have. By having this project done to DBpedia standards, it is ensured that the baseline will be stable and the results will be repeated regularly.
Thanks,
      GerardM

Tuesday, August 27, 2013

The tragedy in a #DBpedia announcement

There was an announcement of achieved goals in the Google Summer of Code for a DBpedia project.
Wikidata integration inside DBpedia we are happy to announce that an initial RDF DBpedia Dumps for Wikidata Data is now available.
Enough reason to read again what this integration is about. The sad thing is in the Additional goals section; "Data quality assurance and tests to improve data quality for academia and enterprises (not for wikipedia infoboxes) would be nice".

The improved data quality is not intended to improve the quality of the data in the Wikipedia infoboxes. This is really sad because as the cooperation with the Deutsche National Bibliothek and the German Wikipedia proves, improvement is beneficial to both parties.

My understanding of why the Wikipedia infoboxes are out of scope is that there has to be a willingness to accept outside 'interference" and we have not learned to cooperate, not even among Wikipedia communities. Wikidata however is the new community and, it has a much higher standard to comply with. The data has to be superior to the data in any Wikipedia because it aims for its data to be used on any Wikipedia.

The best thing to do is side step this issue and make sure that all the questionable data in Wikidata is flagged so that we can start to find out what is best after all.
Thanks,
       GerardM

Sunday, March 13, 2011

The power of nice

A reader of the #Malayalam #Wikipedia, a visually impaired reader at that, send a thank you e-mail to Santhosh because he could follow the text thanks to the Dhvani text to speech application. 

Dhvani is capable of producing intelligible text for 11 Indian languages.
  • Bengali
  • Gujarati
  • Hindi
  • Kannada
  • Malayalam
  • Marathi
  • Oriya
  • Panjabi
  • Tamil
  • Telugu
  • Pashto (experimental)
All these languages have their Wikipedia and it is a happy surprise that a solution for accessibility issues for these Indian languages is a reality.  The Wikipedia reader was grateful for what he had and asked if it would be possible to have  Dhvani express input from the keyboard so that they are helped from "login" to "logout".

As the Silpa project has applied for this years Google Summer of Code, when Silpa is selected that could make this a reality this year.
Thanks,
       GerardM

Sunday, March 28, 2010

The integration of #OpenStreetMap and #MediaWiki needs project management

One of the most eagerly awaited software projects is the integration of OSM in Wikipedia. Having support for high quality maps will provide such an obvious improvement it is actually a no-brainer. When you consider how much effort has gone into bringing maps to Wikipedia; the German chapter investing in it, the time of at least two developers over the years and maps integration as Google Summer of Code projects ...

नीरज अग्रवाल
This year, another student wants to make his mark on the map of the Google Summer of Code. Neeraj Agarwal. Neeraj is from India and if there is one country that will poses a challenge to get all its villages, towns and cities on the Wikipedia map, it is India. Not only because of there being so many of them but also because of there being so many Wikipedias in the languages of the Indian subcontinent.

I had a word with several people interested and involved in the Wikipedia and OpenStreetMap integration and for me it is clear that many people want this but that all the effort is not managed into one cohesive effort. An effort that brings the many interested parties together, an effort that priorities what needs to be done. An effort that coordinates what it takes to get a product out.

At this stage we need the effort of Neeraj, we need our bugmeister to assess the existing code, we need to get an initial product out. Once we have it out, we can make it pretty.
Thanks,
      GerardM

Tuesday, March 13, 2007

Google Summer of Code

Google has announced its third Google Summer of Code. This is an annual event where students develop on Open Source projects. This is definetly one of those activities that does a lot of good. It is one way whereby Google makes its mantra of "do no evil" work well.

For Open Progress, we have entered for a first time; we have a nice mix of MediaWiki and OmegaWiki based projects. All these projects are dear to us. We have shown Brion our list, and we are likely to work together on these.

What struck me is that when you apply for the GSOC, it is compulsory to have a mailing list. This is the traditional way of doing things. I am subscribed to many mailing lists. I think mailing lists suck big-time. There is so much repetition, the signal to noise ration is typically quite bad. I do not understand why people do not use a wiki to document and discuss.

I think this is one of those instances where software development proves to be conservative. When you follow the subjects you are interested in on a wiki, you can use watch lists to make a selection, you can use RSS to follow the changes on a low bandwidth wiki.

Because you refactor what is there, there is no need to repeat so much. When you have discussions that are getting out of hand, backrooms can be opened for those quarrelling. Maybe I am an idealist that I see it in this way .. oh well ..

Thanks,
GerardM