ELAG 2014
Discovering libraries’ gold through collections-level descriptions
Valentine Charles, Data Specialist at The European Library and Europeana
(See also description of talk)
Europeana work in a large scale aggregation ecosystem and now works in collaboration with Cendari. Digitisation is still only the tip of the iceberg of the content of European libraries. Digital objects are displayed attractively as they are very visual. There is also full text available. But most of this material is disconnected from one another. Mostly it is item-level description, with different levels of quality. Wouldn't it be nice to link an image to the relevant journal page?
New strategy for collection-level description. Looking at specific topics, talking to historians, surveying members. E.g. Old Slavic Manuscripts - not digitised but at least there is some description. Another example would be not so much a subject but the specific collections from a particular library, e.g. National Library of Serbia.
Collaboration with Cendari with libraries and archives, about integrating digital data on the mediaval times and the First World War. Researchers working on a project should use this data and we try to facilitate the research activity.The aim is to link the material from different libraries. An environment called Archival research guide is being built to support research. Idea is: when a researcher starts on a topic, he/she writes some paragraphs and it would be incorporated in the guide so that it becomes a narrative, which links to specific collections (the sources). Encouraged to use and re-use data directly in the research, rather than only talking about it. Tools such as NER (name-entity recognition) technology to help identify the entities used in the research or facets etc. are made available. The Archival research guides provide access points to relevant contemporary research, connect collection description to other resources via domain specific ontologies.
Beyond collection description, the most interesting is to link it to other data and other type of material. E.g. with some full text, enriching with annotations, vocabularies, NER etc. Also interested to get this data re-integrated in the various libraries. Cendari is a 4-year project, there are 2 more years to go.
Showing posts with label metadata. Show all posts
Showing posts with label metadata. Show all posts
Thursday, 12 June 2014
Wednesday, 11 June 2014
MIF and Europeana Inside
ELAG 2014
Metadata Interoperability Framework (MIF)
Naeem Muhammad, Software Architect at LIBIS KULeuven and Sam Alloing, Business Consultant at LIBIS KULeuven, Belgium
Made for Europeana inside. This is a technical project with different partners, content providers and software providers. It's to create a better integration of the different content in Europeana. So the goal is to create a component that developers of the different systems can directly add into their content management system and it will talk to Europeana. End of this project planned in Sept. 2014.
The content is enriched by Europeana and the content providers can get it back. This is still in discussion and development. The enriched metadata is not always correct so there are still issues to resolve.
ECK=Europeana Connection Kit. Technical providers use this to transform and push data to Europeana. The ECK local is the part to be integrated in the local system itself. Core ECK services include:
Mapping and Transformation supports MARC to EDM and LIDO to EDM because that's what used in Europeana. LIDO =XML fomrat used by museums, EDM=RDF format from Europeana. So the input has to be MARC XML or LIDO. The output is only EDM at the moment. There are core classes and additional classes. EDM uses Dublin Core. For MARC, it works like this:
[command],[marc tag + subfield],[edm field], e.g. COPY, marc506a,dc:rights
doesn't use indicators at this moment, it could change.
Commands are: COPY, APPEND, SPLIT, COMBINE (multiple source fields can be combined in one target field), LIMIT (to limit the number of characters in a field), PUT, REPLACE, CONDITION (combine different actions and use a conditional flow; can be used with IF.
The plan is that EDM has to be as easy for users as possible, even though some understanding will help. The important is to know the EDM field, not the format.
It's a wservice, so no user interface. Meant to be integrated in CMS or use a REST client. Parameters: records (can be a zip file, XML), mappingRulesFile, sourceFormat (LIDO or MARC), targetFormat (EDM)
Future: add input formats, such as csv, filemaker xml, some custom xml etc. Update/add output formats (add EDM contextual classes, other formats...), extend/add update actions, add queuing (near future), add mapping interface (or integrate with MINT, another Europeana project for mapping)
Metadata Interoperability Framework (MIF)
Naeem Muhammad, Software Architect at LIBIS KULeuven and Sam Alloing, Business Consultant at LIBIS KULeuven, Belgium
Made for Europeana inside. This is a technical project with different partners, content providers and software providers. It's to create a better integration of the different content in Europeana. So the goal is to create a component that developers of the different systems can directly add into their content management system and it will talk to Europeana. End of this project planned in Sept. 2014.
The content is enriched by Europeana and the content providers can get it back. This is still in discussion and development. The enriched metadata is not always correct so there are still issues to resolve.
ECK=Europeana Connection Kit. Technical providers use this to transform and push data to Europeana. The ECK local is the part to be integrated in the local system itself. Core ECK services include:
- Metadata definintion
- Set Mangaer
- Statistics
- PID generation
- Preview service
- Validation of metadata
- Data push (Sword) / data pull (OAI-PMH) because some content providers would rather push the data to Europeana rather than them taking it but Europeana is afraid of compabtibility issues so the data pull is still the one in use
- Mapping and Transformation
Mapping and Transformation supports MARC to EDM and LIDO to EDM because that's what used in Europeana. LIDO =XML fomrat used by museums, EDM=RDF format from Europeana. So the input has to be MARC XML or LIDO. The output is only EDM at the moment. There are core classes and additional classes. EDM uses Dublin Core. For MARC, it works like this:
[command],[marc tag + subfield],[edm field], e.g. COPY, marc506a,dc:rights
doesn't use indicators at this moment, it could change.
Commands are: COPY, APPEND, SPLIT, COMBINE (multiple source fields can be combined in one target field), LIMIT (to limit the number of characters in a field), PUT, REPLACE, CONDITION (combine different actions and use a conditional flow; can be used with IF.
The plan is that EDM has to be as easy for users as possible, even though some understanding will help. The important is to know the EDM field, not the format.
It's a wservice, so no user interface. Meant to be integrated in CMS or use a REST client. Parameters: records (can be a zip file, XML), mappingRulesFile, sourceFormat (LIDO or MARC), targetFormat (EDM)
Future: add input formats, such as csv, filemaker xml, some custom xml etc. Update/add output formats (add EDM contextual classes, other formats...), extend/add update actions, add queuing (near future), add mapping interface (or integrate with MINT, another Europeana project for mapping)
Details, links and interface
ELAG 2014
The LIBRIS upgrade
Niklas Lindström, Lina Westerling, Swedish National Libraray
Abstract: Starting in earnest in 2012, The Swedish National Library (Kungliga Biblioteket – KB) begun the development of a new infrastructure and system, based at its core on Linked Data. It directly employs the linked entity description model represented by RDF, and has the capacity to mesh with other linked data on the web, through minimal engineering efforts.
(see more description of talk)
Need for a modern produce, a platform for data for making is searchable and describing, a method for mapping exisitng data to contemporary models of description and a user interface for editing (cataloguing, curating, linking). Web-based cataloguing tool.
The platform is Open Source and works with all data formats, including RDF etc. Limits of MARC, especially hard to find things and to define things. RDF is not a solution but a means to help solve this problem, because of how it describes data. The tool is a simple expression independent of formats, terms etc. Transform of MARC in JSON-LD. Use of prefixes and uri's.
The Utter Denormalisation of turning JSON-LD back into MARC. This is a temporary measure, because needs to integrate with union catalogues, extract data etc. But the idea is that there's a new interface and new formats.
The design is intiuitve, simple, inspiring, user centered. See beta at devkat.libris.kb.se (test/test) It is quite similar to an end-user search tool. It is based on linked data. Needs to handle all data. Normalising the catalogue will not be able to cover everything.
Doing the mapping is challenging, the data expressed in MARC isn't always normalised so it's not clear if the description is an expression or a manifestation. MARC is very structured but sometimes meaningless, there's lots of convolusion in the specificity, the perspective of different domaines are not well coordinated etc. But there are possibilities of capturing the specificity, better coordinating the vocabularies and so on. Then by linking to external resources we add more value to our resources. Use of SPARQL to help in the linking of sources. Value can also be added to link to internal data.
Another of the main challenges is convincing people, especially cataloguers, so we need to be open.
The LIBRIS upgrade
Niklas Lindström, Lina Westerling, Swedish National Libraray
Abstract: Starting in earnest in 2012, The Swedish National Library (Kungliga Biblioteket – KB) begun the development of a new infrastructure and system, based at its core on Linked Data. It directly employs the linked entity description model represented by RDF, and has the capacity to mesh with other linked data on the web, through minimal engineering efforts.
(see more description of talk)
Need for a modern produce, a platform for data for making is searchable and describing, a method for mapping exisitng data to contemporary models of description and a user interface for editing (cataloguing, curating, linking). Web-based cataloguing tool.
The platform is Open Source and works with all data formats, including RDF etc. Limits of MARC, especially hard to find things and to define things. RDF is not a solution but a means to help solve this problem, because of how it describes data. The tool is a simple expression independent of formats, terms etc. Transform of MARC in JSON-LD. Use of prefixes and uri's.
The Utter Denormalisation of turning JSON-LD back into MARC. This is a temporary measure, because needs to integrate with union catalogues, extract data etc. But the idea is that there's a new interface and new formats.
The design is intiuitve, simple, inspiring, user centered. See beta at devkat.libris.kb.se (test/test) It is quite similar to an end-user search tool. It is based on linked data. Needs to handle all data. Normalising the catalogue will not be able to cover everything.
Doing the mapping is challenging, the data expressed in MARC isn't always normalised so it's not clear if the description is an expression or a manifestation. MARC is very structured but sometimes meaningless, there's lots of convolusion in the specificity, the perspective of different domaines are not well coordinated etc. But there are possibilities of capturing the specificity, better coordinating the vocabularies and so on. Then by linking to external resources we add more value to our resources. Use of SPARQL to help in the linking of sources. Value can also be added to link to internal data.
Another of the main challenges is convincing people, especially cataloguers, so we need to be open.
Friday, 21 September 2012
Second Linked Open Data Conference
I have been live blogging the Second Linked Open Data Conference organised by CILIP Cataloguing and Indexing Group, which was held on 19th September 2012 in Edinburgh.
This was a very interesting day covering a diversity of approaches to Linked Open Data and the live blog can be found on the CIG's website.
There is also a short url to view this event's live blog: http://bit.ly/PNJSuc
This was a very interesting day covering a diversity of approaches to Linked Open Data and the live blog can be found on the CIG's website.
There is also a short url to view this event's live blog: http://bit.ly/PNJSuc
Labels:
CIG,
CILIP,
conference,
linked data,
live blog,
live blogging,
LOD,
metadata,
open data,
scotland
Thursday, 13 September 2012
IGeLU 2012 - Plenary closing session
Bibliographic Framework Initiative Approach for MARC Data as Linked Data
Sally McCallum, Chief, Network Development and Standards Office, Library of Congress
MARC
Although MARC is 40 years old, it still dominates the environment. There are lots of sharing options on a MARC format based record. It has adjusted to various cataloguing norms. It has lots of data elements even compared to other norms that may be more sophisticated in other respects. MARC has adapted to technical change. There are structural limitations (for example when extending MARC in xml) so we need to move ahead.
RDA and more
There are new cataloguing norms, in particular RDA, but there are others too. Within the RDA ground there is more option for parsing data. It's a 2-way sword because that creates more data elements. There is a use of codes rather than terms and an emphasis on relationships. RDA also offers more flexibility with authoritative headings. Is it possible to include the broader cultural community in library cataloguing norms? We say that and we'll be able to accomodate all the various cultural environments but it is not clear yet that we'll be able to.
Transcriptions
There are pros and cons to transcriptions. As resources are published in more than one way, that is transcribed in more than one way, this is becoming less of something that we have to be focussing on. In the cataloguing area, and headings versus terms, what should we use? At the LoC we use headings, but at we don't know what the future will be. There is also more user supplied information (crowd sourcing).
Type of resources
The printed resource production doesn't seem to go down whilst e-resources is increasing from the publishers as well as in collections. We'll be in a situation where the collection of printed resources is changing. Then there are casual resources, for example, twitter etc. We don't really know what to do with that, should we archive it?
Systems
There is more need for e-resources access management and this should take into account licensing and rights management. E-resource object management implies preservation. There is a lot of push on retrieval needs, both basic and scholar. Libraries have a role to play still in this area.
So the main issue is flexibility. In the next 5 years all of this will have changed again.
Framework Initiative - the bold venture
We need to work together to share bibliographic description and save money. We've included people with broad perspectives and have defined the requirements and the approach for the Initiative.
Requirements:
- Broad accommodation of content norms and data models
- New views of different types of metadata: descriptive, authority, holdings / coded data, classification data, subject data / preservation, rights, technical, archival
- Reconsideration of the activity relationships: exchange, internal storage, inupt interfaces and techniques
- Enhanced linking: traditional = textual, identifiers / semantic technology = URIs
- Accommodate different types of libraries: large, small, research, public, specialised...
- MARC compatibility: maintenance of MARC21 continued / enable reuse of data from MARC / provision of transformations to new models
Approach
Orientation towards the web and linked data. Investigate the use of semantic web standards (RDF data model, various syntaxes: xml, json, n-triples etc.) We want to work with high models and collaboration.
Linked data is important because of the amount of social media on the web, the way search engines work, more and more applications are going towards linked data and there is an increased flexibility to describe resources.
Initial model development
We have a contract with Zepheira (May 2012), because we wanted someone who wasn't completely absorbed in MARC, RDA etc. and who had a broad approach, with a good understanding of DC (?)technologies. So our partner has a long experience of MARC, as well asW3C and a pratcial application of RDF. We had 2 major tasks: Review of several related initiatives and translating bibliographic data to a
linked data form (evolution, not revolution / a basis for cummunity discussion and dialogue).
Balancing factors
- MARC21 historical data and roles
- Previous efforts for modelling bibliographinc information (FRBR - RDA, Indecs - Onix)
- Previous efforts to express bibliographic information as linked data (BL, Deutsche Nazional Bibliothek, Library of Congress' ID, OCLC Worldcat, schema.org)
- Using the web as model for expressing and connecting information (URIs, decentralisation of data, annotation)
- Library community social and technical deployment probabilities
- Adoption outside the library community
- Flexibility for future cataloguing and use scenarios
- Leverage machine technology for the mechanical while keeping the librarian expertise in control
So we started by descontructing MARC, that means identifying MARC resources (MARCR), for example people, places, institutions, subjects etc. Since you need a replacement for those, they have to be pulled out first.
Phase 1 - High level model
4 core classes (initially 2 but rare books / music librarians felt this was not enough):
- Work: resource reflecting the conceptual essence of the cataloguing item / roughly equivalent to FRBR work or expression
- Instance: resource reflecting an individual, material embodiment of the Work
- Authority: resource reflecting key authority concepts that have defined relationships
- Annotation: resource that "decorates" other MARCR resources (e.g. holdings, cover images, reviews)
Each of these are represented by URIs.
So subjects and creators relate to Work; publisher and format relate to Instance; Instance and Work are linked together. Annotations can relate to either Work or Instance so there are lots of places that we can use URIs.See model photos here
Phase 1.5 - early experimentation
- Preliminary work at the LoC
- Very small group of early experimenters
- Working with high level model, vocabularies, conversion tools
- Creative development of syntaxes and configurations
- Adjust model
Model development
- Make model, mappings and tools available and encourage broader experimentation?
- Parallel phase 2 to refine the model and keep folding in experience based changes
- Follow the progress: www.loc.gov/marc/transition
- Join the discussion: bibframe@listserv.loc.gov
Sally McCallum, Chief, Network Development and Standards Office, Library of Congress
MARC
Although MARC is 40 years old, it still dominates the environment. There are lots of sharing options on a MARC format based record. It has adjusted to various cataloguing norms. It has lots of data elements even compared to other norms that may be more sophisticated in other respects. MARC has adapted to technical change. There are structural limitations (for example when extending MARC in xml) so we need to move ahead.
RDA and more
There are new cataloguing norms, in particular RDA, but there are others too. Within the RDA ground there is more option for parsing data. It's a 2-way sword because that creates more data elements. There is a use of codes rather than terms and an emphasis on relationships. RDA also offers more flexibility with authoritative headings. Is it possible to include the broader cultural community in library cataloguing norms? We say that and we'll be able to accomodate all the various cultural environments but it is not clear yet that we'll be able to.
Transcriptions
There are pros and cons to transcriptions. As resources are published in more than one way, that is transcribed in more than one way, this is becoming less of something that we have to be focussing on. In the cataloguing area, and headings versus terms, what should we use? At the LoC we use headings, but at we don't know what the future will be. There is also more user supplied information (crowd sourcing).
Type of resources
The printed resource production doesn't seem to go down whilst e-resources is increasing from the publishers as well as in collections. We'll be in a situation where the collection of printed resources is changing. Then there are casual resources, for example, twitter etc. We don't really know what to do with that, should we archive it?
Systems
There is more need for e-resources access management and this should take into account licensing and rights management. E-resource object management implies preservation. There is a lot of push on retrieval needs, both basic and scholar. Libraries have a role to play still in this area.
So the main issue is flexibility. In the next 5 years all of this will have changed again.
Framework Initiative - the bold venture
We need to work together to share bibliographic description and save money. We've included people with broad perspectives and have defined the requirements and the approach for the Initiative.
Requirements:
- Broad accommodation of content norms and data models
- New views of different types of metadata: descriptive, authority, holdings / coded data, classification data, subject data / preservation, rights, technical, archival
- Reconsideration of the activity relationships: exchange, internal storage, inupt interfaces and techniques
- Enhanced linking: traditional = textual, identifiers / semantic technology = URIs
- Accommodate different types of libraries: large, small, research, public, specialised...
- MARC compatibility: maintenance of MARC21 continued / enable reuse of data from MARC / provision of transformations to new models
Approach
Orientation towards the web and linked data. Investigate the use of semantic web standards (RDF data model, various syntaxes: xml, json, n-triples etc.) We want to work with high models and collaboration.
Linked data is important because of the amount of social media on the web, the way search engines work, more and more applications are going towards linked data and there is an increased flexibility to describe resources.
Initial model development
We have a contract with Zepheira (May 2012), because we wanted someone who wasn't completely absorbed in MARC, RDA etc. and who had a broad approach, with a good understanding of DC (?)technologies. So our partner has a long experience of MARC, as well asW3C and a pratcial application of RDF. We had 2 major tasks: Review of several related initiatives and translating bibliographic data to a
linked data form (evolution, not revolution / a basis for cummunity discussion and dialogue).
Balancing factors
- MARC21 historical data and roles
- Previous efforts for modelling bibliographinc information (FRBR - RDA, Indecs - Onix)
- Previous efforts to express bibliographic information as linked data (BL, Deutsche Nazional Bibliothek, Library of Congress' ID, OCLC Worldcat, schema.org)
- Using the web as model for expressing and connecting information (URIs, decentralisation of data, annotation)
- Library community social and technical deployment probabilities
- Adoption outside the library community
- Flexibility for future cataloguing and use scenarios
- Leverage machine technology for the mechanical while keeping the librarian expertise in control
So we started by descontructing MARC, that means identifying MARC resources (MARCR), for example people, places, institutions, subjects etc. Since you need a replacement for those, they have to be pulled out first.
Phase 1 - High level model
4 core classes (initially 2 but rare books / music librarians felt this was not enough):
- Work: resource reflecting the conceptual essence of the cataloguing item / roughly equivalent to FRBR work or expression
- Instance: resource reflecting an individual, material embodiment of the Work
- Authority: resource reflecting key authority concepts that have defined relationships
- Annotation: resource that "decorates" other MARCR resources (e.g. holdings, cover images, reviews)
Each of these are represented by URIs.
So subjects and creators relate to Work; publisher and format relate to Instance; Instance and Work are linked together. Annotations can relate to either Work or Instance so there are lots of places that we can use URIs.See model photos here
Phase 1.5 - early experimentation
- Preliminary work at the LoC
- Very small group of early experimenters
- Working with high level model, vocabularies, conversion tools
- Creative development of syntaxes and configurations
- Adjust model
Model development
- Make model, mappings and tools available and encourage broader experimentation?
- Parallel phase 2 to refine the model and keep folding in experience based changes
- Follow the progress: www.loc.gov/marc/transition
- Join the discussion: bibframe@listserv.loc.gov
Subscribe to:
Posts (Atom)