You are browsing the archive for wp3.

Community Discussions 3

- July 13, 2012 in BibServer, Data, event, Events, JISC OpenBib, jiscopenbib2, licensing, News, OKFN Openbiblio, wp3, wp4, wp5

It has been a couple of months since the round-up on Community Discussions 2 and we have been busy! BiblioHack was a highlight for me, and last week included a meeting of many OKFN types – here’s a picture taken by Lucy Chambers for @OKFN of some team members: IMG_0351 The Discussion List has been busy too:
  • Further to David Weinbergers’s pointer that Harvard released 12 million bibliographic records with a CC0 licence, Rufus Pollock created a collection on the DataHub and added it to the Biblio section for easy of reference

  • Rufus also noticed that OCLC had issued their major release of VIAF, meaning that millions of author records are now available as Open Data (under Open Data Commons Attribution license), and updated the DataHub dataset to reflect this

  • Peter Murray-Rust noted that Nature has made its metadata Open CC0

  • David Shotton promoted the International Workshop on Contributorship and Scholarly Attribution at Harvard, and prepared a handy guide for attribution of submissions

  • Adrian Pohl circulated a call for participation for the SWIB12 “Semantic Web in Bibliotheken” (Semantic Web in Libraries) Conference in Cologne, 26-28 November this year, and hosted the monthly Working Group call

  • Lars Aronsson looked at multivolume works, asking whether the OpenLibrary can create and connect records for each volume. HathiTrust and Gallica were suggested as potential tools in collating volumes, and the barcode (containing information populated by the source library) was noted as being invaluable in processing these

  • Sam Leon explained that TEXTUS would be integrating BibSever facet view and encouraged people to have a look at the work so far; Tom Oinn highlighted the collaboration between Enriched BibJSON and TEXTUS, and explained that he would be adding a ‘TEXTUS’ field to BibJSON for this purpose

  • Sam also circulated two tools for people to test, Pundit and Korbo, which have been developed out of Digitised Manuscripts to Europeana (DM2E)

  • Jenny Molloy promoted the Open Science Hackday which took place last week – see below for a snap-shot courtesy of @OKFN:

IMG_1964 In related news, Peter Murray-Rust is continuing to advocate the cause of open data – do have a read of the latest posts on his blog to see how he’s getting on. The Open Biblio community continues to be invaluable to the Open GLAM, Heritage, Access and other groups too and I would encourage those interested in such discussions to join up at the OKFN Lists page.

BiblioHack: Day 2, part 2

- June 14, 2012 in BibServer, Data, event, Events, JISC OpenBib, jiscopenbib2, minutes, News, OKFN Openbiblio, Talks, wp1, wp2, wp3, wp4, wp5, wp6, wp7, wp8, wp9

Pens down! Or, rather, key-strokes cease! BiblioHack has drawn to a close and the results of two days’ hard labour are in:

A Bibliographic Toolkit

Utilising BibServer Peter Murray-Rust reported back on what was planned, what was done, and the overlap between the two! The priority was cleaning up the process for setting up BibServers and getting them running on different architectures. (PubCrawler was going to be run on BibServer but currently it’s not working). Yesterday’s big news was that Nature has released 30 million references or thereabouts – this furthers the cause of scholarly literature whereby we, in principle, can index records rather than just corporate organisations being able / permitted to do so. National Bibliographies have been put on BibSoup – UK (‘BL’), Germany, Spain and Sweden – with the technical problem character encodings raising its head (UTF8 solves this where used). Also, BibSoup is useful for TEXTUS so the overall ‘toolkit’ approach is reinforced! Open Access Index Emanuil Tolev presented on ACat – Academic Catalogue. The first part of an index is having things to access – so gathering about 55,000 journals was a good start! Using Elastic Search within these journals will give list of contents which will then provide lists of articles (via facet view), then other services will determine licensing / open access information (URL checks assisted in this process). The ongoing plan is to use this tool to ascertain licensing information for every single record in the world. (Link to ACat to follow). Annotation Tools Tom Oinn talked about the ideas that have come out of discussions and hacking around annotators and TEXTUS. Reading lists and citation management is a key part of what TEXTUS is intended to assist with, so the plan is for any annotation to be allowed to carry a citation – whether personal opinion or related record. Personalised lists will come out of this and TEXTUS should become a reference management tool in its own right. Keep your eye on TEXTUS for the practical applications of these ideas! Note: more detailed write-ups will appear courtesy of others, do watch the OKFN blog for this and all things open… Postscript: OKFN blog post here Huge thanks to all those who participated in the event – your ideas and enthusiasm have made this so much fun to be involved with. Also thanks to those who helped run the event, visible or behind-the-scenes, particularly Sam Leon. Here’s to the next one :-)

BiblioHack: Day 2, part 1

- June 14, 2012 in BibServer, Data, event, Events, JISC OpenBib, jiscopenbib2, minutes, News, OKFN Openbiblio, Talks, wp1, wp2, wp3, wp4, wp5, wp6, wp7, wp8, wp9

After easing into the day with breakfast and coffee, each of the 3 sub-groups gave an overview of the mini-project’s aim and fed back on the evening’s progress:
  • Peter Murray-Rust revisited the overarching theme of ‘A Bibliographic Toolkit’ and the BibServer sub-group’s specific work on adding datasets and easily deploying BibServer; Adrian Pohl followed up to explain that he would be developing a National Libraries BibServer.
  • Tom Oinn explained the Annotation Tools sub-groups’s work on developing annotation tools – ie TEXTUS – looking at adding fragments of text, with your own comments and metadata linked to it, which then forms BibSoup collections. Collating personalised references is enhanced with existing search functionality, and reading lists with annotations can refer to other texts within TEXTUS.
  • Mark MacGillivray presented the 3rd group’s work on an Open Access Index. This began with listing all the journals that can be found in the whole world, with the aim of identifying the licence of each article. They have been scraping collections (eg PubMed) and gathering journals – at the time of speaking they had around 50,000+! The aim is to enable a crowd-sourced list of every journal in the world which, using PubCrawler, should provide every single article in the world.
With just 5 hours left before stopping to gather thoughts, write-up and feedback to the rest of the group, it will be very interesting to see the result…

BiblioHack: Day 1

- June 14, 2012 in BibServer, Data, event, Events, JISC OpenBib, jiscopenbib2, licensing, lod-lam, minutes, OKFN Openbiblio, Talks, wp1, wp2, wp3, wp4, wp5, wp6, wp7, wp8, wp9

The first day of BiblioHack was a day of combinations and sub-divisions! The event attendees started the day all together, both hackers and workshop / seminar attendees, and Sam introduced the purpose of the day as follows: coders – to build tools and share ideas about things that will make our shared cultural heritage and knowledge commons more accessible and useful; non-coders – to get a crash course in what openness means for galleries, libraries, archives and museums, why it’s important and how you can begin opening up your data; everyone – to get a better idea about what other people working in your domain do and engender a better understanding between librarians, academics, curators, artists and technologists, in order to foster the creation of better, cooler tools that respond to the needs of our communities. The hackers began the day with an overview of what a hackathon is for and how it can be run, as presented by Mahendra Mahey, and followed with lightning talks as follows:
  • Talk 1 Peter Murray Rust & Ross Mounce – Content and Data Mining and a PDF extractor
  • Talk 2 Mike Jones – the m-biblio project
  • Talk 4 Ian Stuart – ORI/RJB (formerly OA-RJ)
  • Talk 5 Etienne Posthumus – Making a BibServer Parser
  • Talk 6 Emanuil Tolev – IDFind – identifying identifiers (“Feedback and real user needs won’t gather themselves”)
  • Talk 7 Mark MacGillivray – BibServer – what the project has been doing recently, how that ties into the open access index idea.
  • Talk 8 Tom Oinn – TEXTUS
  • Talk 9 Simone Fonda – Pundit – collaborative semantic annotations of texts (Semantic Web-related tool)
  • Talk 10 Ian Stuart – The basics of Linked Data
We decided we wanted to work as a community, using our different skills towards one overarching goal, rather than breaking into smaller groups with separate agendas. We formed the central idea of an ‘open bibliographic tool-kit’ and people identified three main areas to hack around, playing to their skills and interests:
  • Utilising BibServer – adding datasets and using PubCrawler
  • Creating an Open Access Index
  • Developing annotation tools
At this point we all broke for lunch, and the workshoppers and hackers mingled together. As hoped, conversations sprung up between people from the two different groups and it was great to see suggestions arising from shared ideas and applications of one group being explained to the theories of the other. We re-grouped and the workshop continued until 16.00 – see here for Tim Hodson’s excellent write-up of the event and talks given – when the hackers were joined by some who attended the workshop. Each group gave a quick update on status, to try to persuade the new additions to the group to join their particular work-flow, and each group grew in number. After more hushed discussions and typing, the day finished with a talk from Tara Taubman about her background in the legalities of online security and IP, and we went for dinner. Hacking continued afterwards and we celebrated a hard day’s work down the pub, lookong forward to what was to come. Day 2 to follow…

BiblioHack Meet-up

- June 13, 2012 in event, Events, JISC OpenBib, jiscopenbib2, OKFN Openbiblio, wp3, wp4

I’ve been quiet on this blog lately, but it’s in the same way a duck looks still when swimming: things may look peaceful but there is much activity going on beneath the surface! The Open Biblio crowd have been busy on the discussion List (link to follow) and the BiblioHack organisers have been preparing for this week’s events, which kicked off with a Meet-up last night. The pre-BiblioHack Meet-up was designed to be an informal opportunity for those involved in the events to put names to faces and start up discussions; it was also open to anyone who wanted to come along to find out more about open data and the OKFN’s Working Groups including Open GLAM, and projects such as DM2E as well as Open Biblio. With no formal agenda, we started up conversations as the mood took us – this covered legalities of openness in relation to IP, licensing and open access, annotation, cat-sitting and the Blues. In a nod to the more ‘usual’ OKFN #OpenData meet-ups, we went around the room to introduce ourselves (trying to explain our interests in only 3 words was challenging…) which prompted some people to cross the room in a purposeful fashion to intercept someone they hadn’t spoken to by that point. I really enjoyed meeting the people with whom I’d be spending the next two days, so thanks to all those who came along, for their interesting ideas and suggestions, and huge thanks to Sam Leon for arranging the tasty food and drinks at C4CC and for facilitating the evening.

Open source development – how we are doing

- May 29, 2012 in BibServer, JISC OpenBib, jiscopenbib2, licensing, progress, progressPosts, projectMethodology, projectPlan, riskAnalysis, software, WIN, wp10, wp2, wp3, wp6, wp9

Whilst at Open Source Junction earlier this year, I talked to Sander van der Waal and Rowan Wilson about the problems of doing open source development. Sander and Rowan work at OSS watch, and their aim is to make sure that open source software development delivers its potential to UK HEI and research; so, I thought it would be good to get their feedback on how our project is doing, and if there is anything we are getting wrong or could improve on. It struck me that as other JISC projects such as ours are required to make their output similarly publicly available, this discussion may be of benefit to others; after all, not everyone knows what open source software is, let alone the complexities that can arise from trying to create such software. Whilst we cannot help avoid all such complexities, we can at least detail what we have found helpful to date, and how OSS Watch view our efforts. I provided Sander and Rowan a review of our project, and Rowan provided some feedback confirming that overall we are doing a good job, although we lack a listing of the other open source software our project relies on, and their licenses. Whilst such data can be discerned from the dependencies of the project, this is not clear enough; I will add a written list of dependencies to the README. The response we received is provided below, followed by the overview I initially provided, which gives a brief overview of how we managed our open source development efforts: ==== Rowan Wilson, OSS Watch, responds: Your work on this project is extremely impressive. You have the systems in place that we recommend for open development and creation of community around software, and you are using them. As an outsider I am able to quickly see that your project is active and the mailing list and roadmap present information about ways in which I could participate. One thing I could not find, although this may be my fault, is a list of third party software within the distribution. This may well be because there is none, but it’s something I would generally be keen to see for the purposes of auditing licence compatibility. Overall though I commend you on how tangible and visible the development work on this project is, and on the focus on user-base expansion that is evident on the mailing list. ==== Mark MacGillivray wrote: Background – May 2011, OKF / AIM bibserver project Open Knowledge Foundation contracted with American Institute of Mathematics under the direction of Jim Pitman in the dept. of Maths and Stats at UC Berkeley. The purpose of the project was to create an open source software repository named BibServer, and to develop a software tool that could be deployed by anyone requiring an easy way to put and share bibliographic records online. A repository was created at http://github.com/okfn/bibserver, and it performs the usual logging of commits and other activities expected of a modern DVCS system. This work was completed in September 2011, and the repository has been available since the start of that project with a GNU Affero GPL v3 licence attached. October 2011 – JISC Open Biblio 2 project The JISC Open BIblio 2 project chose to build on the open source software tool named BibServer. As there was no support from AIM for maintaining the BibServer repository, the project took on maintenance of the repository and all further development work, with no change to previous licence conditions. We made this choice as we perceive open source licensing as a benefit rather than a threat; it fit very well with the requirements of JISC and with the desires of the developers involved in the project. At worst, an owner may change the licence attached to some software, but even in such a situation we could continue our work by forking from the last available open source version (presuming that licence conditions cannot be altered retrospectively). The code continues to display the licence under which it is available, and remains publicly downloadable at http://github.com/okfn/bibserver. Should this hosting resource become publicly unavailable, an alternative public host would be sought. Development work and discussion has been managed publicly, via a combination of the project website at http://openbiblio.net/p/jiscopenbib2, the issue tracker at http://github.com/okfn/bibserver/issues, a project wiki at http://wiki.okfn.org/Projects/openbibliography, and via a mailing list at openbiblio-dev@lists.okfn.org February 2012 – JISC Open Biblio 2 offers bibsoup.net beta service In February the JISC Open Biblio 2 project announced a beta service available online for free public use at http://bibsoup.net. The website runs an instance of BibServer, and highlights that the code is open source and available (linking to the repository) to anyone who wishes to use it. Current status We believe that we have made sensible decisions in choosing open source software for our project, and have made all efforts to promote the fact that the code is freely and publicly available. We have found the open source development paradigm to be highly beneficial – it has enabled us to publicly share all the work we have done on the project, increasing engagement with potential users and also with collaborators; we have also been able to take advantage of other open source software during the project, incorporating it into our work to enable faster development and improved outcomes. We continue to develop code for the benefit of people wishing to publicly put and share their bibliographies online, and all our outputs will continue to be publicly available beyond the end of the current project.

BiblioHack hackathon registration form

- May 9, 2012 in event, Events, JISC OpenBib, jiscopenbib2, OKFN Openbiblio, wp3, wp4

To register for BiblioHack, the 2-day hackathon event on 13th-14th June 2012, please submit your details using the form below and ensure you scroll down to complete all fields. Please note, unfortunately spaces are limited so completion of form does not guarantee a place, but we will do our best to accommodate as many people as possible.

Recent BibServer technical development

- May 8, 2012 in BibServer, Data, JISC OpenBib, jiscopenbib2, News, OKFN Openbiblio, wp2, wp3, wp5, wp6, wp7, wp8

Along with the recent push of new front-end functionality to BibServer, and demonstrated on BibSoup, we have also applied some changes to the back-end. The new scheduled collection uploader is now runnable as a stand-alone tool, to which source URLs can be provided for retrieval, conversion, and upload. Retrieved sources are stored and available from a folder on disk, as are the conversions. Parsers can now be written in any language and plugged into the ingest functionality – for example, we now have a MARC parser that runs in perl and is usable via ingest.py and available on an instance of BibServer – thanks very much to Ed for that. In addition, parsers need no longer be ‘parsers’ – we have introduced the concept of scrapers as well. Check out our new Wikipedia parser / scraper, for example; it functions by taking in a search value rather than a URL, then using that to search Wikipedia for relevant references which it downloads, bundles, and converts to a BibJSON collection – this is a really great example that Etienne put together, and it demonstrates a great deal of potential for further parser / scraper development. See the examples on the BibServer repo for more insight – they are in the parserscrapers_plugins folder, and they are managed by bibserver/ingest.py. We know documents are now lacking – we have set up an online docs resource but are in the process of writing up to populate it – please check back soon. As usual, development work is scheduled via the tickets and milestones on our repo. Current efforts are on documentation and adding as many feature requests as possible before our hackathon on June 12th – 14th.

BibJSON updates

- May 8, 2012 in BibServer, Data, JISC OpenBib, jiscopenbib2, lod-lam, News, OKFN Openbiblio, wp2, wp3, wp5, wp6, wp7, wp8

Following recent discussion on our mailing list, BibJSON has been updated to adopt JSON-LD for all your linked data needs. This enables us to keep the core of BibJSON pretty simple whilst also opening up potential for more complex usage where that is required. Due to this, we no longer use the “namespace” key in BibJSON. Other changes include usage of “_” prefix on internal keys – so wherever our own database writes info into a record, we prefix it, such as “_id”. Because of this, uploaded BibJSON records can have an “id” key that will work, as well as an “_id” uuid applied by the BibServer system. For more information, check out BibJSON.org and JSON-LD

New BibServer features available on BibSoup

- May 8, 2012 in BibServer, Data, JISC OpenBib, jiscopenbib2, News, OKFN Openbiblio, wp2, wp3, wp5, wp6, wp7, wp8

A couple of months ago the development team had a Sprint and came up with some cool ideas of how to improve the user experience for BibServer and, subsequently, BibSoup. Have a play with the new features and see below for the details:

Main pages

  • Collections visualisation – a smart new graphic on the landing page showing information from new collections

  • Improved FAQ section with links to videos (coming soon: links to our new online docs)

Creating collections

  • New Wikipedia parser – create a collection based on the references retrievable from Wikipedia for your chosen search value

  • Improved collection upload – specify collection information, then view upload tickets to see progress and errors

  • ‘Retry’ and other options on particular collection creation attempts are also now available from the tickets page

Search results

  • Filter search results by a value range as well as specific values

  • Visualise any filter as a bubble chart and select the values you want to search with

  • Add / remove available filters and rename filter display names

  • Improved layout of record info in search results, including auto-display of the first image referenced in a record – e.g. if there is a link to an image in your record, it is displayed in the search result

Managing and sharing collections

  • Collection admin available – save your current display settings as the default for your collection, allow other users to have admin rights on your own collection

  • Share any specific searches by providing the URL displayed under the ‘share’ option

  • Embed – as the whole front-end of search and collection visualisation is handled by facetview it is possible to embed your collection search in any web page you control; the share / embed option on collection pages provides the code you need to insert to enable this

  • Download as BibJSON – a nice new obvious button on each collection provides a link to download your collection as BibJSON

Viewing records

  • Improved display of individual records, including search options to discover relevant content online

  • EXPERIMENTAL record editing – this has been enabled although still in progress – you can edit the content of a record using a visual display of the keys and values in the record, although functionality for adding new keys does not yet work. However, you can also edit the JSON directly via the options, and try saving that. Be aware – this could damage your records, and of course changes the details from whatever they were in the source content.

Still in development

These ones are not yet available on BibSoup but watch this space:

  • Creating new collections on-site – search and find particular records for inclusion in new collections or addition to pre-existing collections. This is not currently possible but we are working on making this an easy process
  • Merging collections
  • Better user creation and management, plus gravatars
  • Additional functionality on record pages – linking out directly to related sources such as PubMed, Total Impact, Service Core etc
We hope you like these changes, and find them useful – do let us know what you think and keep an eye out for the upcoming improvements.