Pages

Showing posts with label British Library. Show all posts
Showing posts with label British Library. Show all posts

Wednesday, 17 September 2014

IGeLU 2014
Deep search or harvest, how do we decide
Simon Moffatt and Illaria Corda, British Library

Context: increasing content, especially digital from a diversity of sources as well as migration form other systems. So there are two options to integrate all this data: through harvesting or with a deep search.
Harvest: put all the data in the central index (Primo)
Deep search: add a new index, so user searches Primo and this other index in a federated way

The decision that the BL had to make was for two major indexes: one was for the replacement of Content Management System and the other one was for the Web Archive

CMS: harvest

The main reasons for choosing harvesting is that the CMS has it's own index, updated daily. Work involved for deep search would be normalisation rules, setting up daily processes etc. The index is not good enough. It only has 50.000 pages which for Primo is not that much.

Records in the CMS
13m from Aleph
5m Sirsi Dynix for sound archive
48m articles from own article catalogue
1m records from other catalogues

Service challenges for 67m records
The index is a very big file (100GB). Overnight schedules are tight. Fast data updates are not possible. Re-indexing is always at least 3hrs, sometimes more. System restart takes 5hrs, re-synch of the search schema is a whole day and the failover system must be available at all times. Must be careful for Primo service packs and hotfixes. Doing the whole index in one go wouldn't be possible so beware when the documentation says "delete index and recreate".

Development challenges
Major normalisations only 3 times a day, need to be careful about impact on services. Implementing Primo enhancements also needs to be considered. In their development environment they have created a sample of all their different datasets. It's also used to test which version of Primo to use etc. Errors have to be found at this point.

But the compensations of big index are:
Speed
Control over the data
Consistency of rules
Common development skills
Single point of failure

Web Archive : deep search

Figures
1.3m documents (pages)
Regular crawl processes
Index size 3.5 terabytes
80-node hadoop cluster where crawls are processed through submission

Implications of choosing the deep search
On the Primo interface there is a tab for the local content and one for the web archive, but otherwise the GUI is the same as are the type of facets etc. Clicking on a web archive record takes to the Ericom system (digitals rights system for e-legal deposit material).

Service challenges for deep search
Introduction of a different point of failure
New troubleshooting procedures
Changes to the solr schema and could break the search
Primo updates can also potentially break the search

Development challenges
Significant development work, e.g. accommodate Primo features such as multi-select faceting, consistency between local and deep search, etc.

But the compensations are:
Ideal for a large index with frequent updates
Indepth indexing
Maintenance of existing scheduled processes
Active community of api developers

Conclusion: Main criteria to decide

Frequency of updates and lots of records then deep is probably better
Budget to buy servers etc.
Development time and skills expertise
Impact on current processes

Questions
Security layer in development api that doesn't expose the solr index
Limitations of Primo: scalability; would exl make local indexes bigger? BL's primo is not hosted. Before work on primo central, we had genuine problems and we worked on them to resolve them. Primo is actually scalable.

Friday, 13 June 2014

Identifying and identifiers

ELAG 2014
Integrating ORCiD – A two way conversation
Tom Demeranville, software engineer specialising in digital identifiers and identities at the British Library

(see description of the talk)

ODIN (DataCite Interoperability Network) is concerned with linking authors with research output and is a 2-year project. It's also looking at datasets, grey literature, etc. What do we mean by identifying authors? The answer varies. One person can have lots of identifiers and profiles, including institutional profile, an ISNI, a ScopusID etc. or an ORCiD.

So the first distinction is between an identifier and a profile. We usually think of identifiers as unique ID but a profile can be much more. Another important point is that no one wants to type the same thing twice. Profiles can be automated or manual. Then there is the difference of identifiers as Institutions or Users, with conflicting notions of control but we all want disambiguation...

So what we need... One identifier and many profiles that solve different use cases.

ORCiD is meant to be a more open identifying system, managed for people with many different use cases. Relevant to publishers, unis, funders and libraries. It would help systems to talk to each other.

Ethos is e-Thesis Online Import. See demo at http://ethos-orcid.appspot.com

The ODIN project is working to integrate ORCiD and DataCite.

The Mechanical Curator

ELAG 2014
The Surprising Adventures of the Mechanical Curator, and Other Tales
Ben O’Steen, technical leader of  British Library Labs

(see description of the talk and the slides)

This project started last year, as an accident! Taking the stuff that's technically accessible... and making it accessible! Engaged with the researchers, formally and informally through yearly competitions. What they win is our time and effort! The unifying theme to (pretty much) all the requests is: Give us everything! But this is quite depressing: so librarians don't take part in research, they're only there to provide content? Another theme is to have tools to interpret the content, to be able to work on broad sweeps of content rather than one at a time.

The Sample Generator shows the chasm between the collection and the digitised material. Not only is the content not as much digitised but it is also not as accessible as it could be.

The challenge was that research didn't want to work with api's but access large amounts of data. Made an experiment: face detection on 19thC illustrations - it wasn't very successful. The depiction is usually "clean" and posed, males represented differently from females and therefore less often detected etc. But it gave the idea of the mechanical curator, who digs in the collection of digitised images and tries to find visually similar images, based on a calculated match. It has now been doing that for a couple of months (and tweets about it). An unguided way of discovering material.

Images published on Flickr, many views in 4 days. They are published as CC0 and there are already examples of creative re-uses, such as colouring-in for children, an artit's interpretation etc. But this doesn't bring money to the Library, which is always hard to justify. But this is encouraging creativity, it may not be research but it's not less imporatnt. An animation student used images to represent them in 3-D. Moments, by Joe Bell

So the impact is hard to measure. Accessible is great, can we make it more useful? A group of UCL big Data CS students will be given access to all the book data, cloud computing and will make an experiment for broader and more direct access to the collections.


Wednesday, 11 June 2014

Role of libraries in supporting digital scholarship

ELAG 2014
Key note: The Role of libraries in supporting digital scholarship
Stella Wisdom, Digital Curator, The British Library

(see description of the talk)

Need to change the services to meet the need of researchers. There is more and more digital content, increased collaboration working or re-purposing of content. The BL wants the researchers to do innovative research with their content. The BL has been digitising for at least two decades and aims to do much more.

Digital content - examples of recent developments at the BL

  • Georeference maps and new interactive tool (http://www.bl.uk/maps
  • Europeana 1914-1918 Roadshows - visited museums in different parts of the country, showing some of their digitised images
  • Off the Map: video games festival, following a preservation about complexe object conference to which Stella went to and gave her ideas of what the BL can do in this area. There's a museum about video games, Victoria & Albert Museum also organised a competition. The BL made a special feature on the web archive. BL organised a competion: Crytek off the map: visual trip through 17thC. London made by 6 2nd-grade students (winners of last year's competition)
  • Work done on sound collections, with permissions to re-use (under certain conditions) see the Flying Buttress
  • Organising exhibitions such as Beautiful Science (picturing data, inspiring insight)
  • Dora's lost data game
  • British Library labs: one of the main actions is a yearly competition to identify innovative ideas that showcase the Library's collections
  • The Victorian Meme Machine, to preserve Victorian jokes (one of the winners of the labs competition) - it will combine jokes with images, all coming from the BL collections