Monday, September 15, 2008

Summer 2008 Project

This summer I worked on a proof of concept project for creating a digital collection of the University of Iowa student newspaper dating from 1869 to the present. The scanning of the newspapers was outsourced to a local company with the product being TIFF files.

To begin the project I researched practices of other similar digital collections and the best practices for newspaper digitization. Some of the examples that I looked at were:
Utah Digital Newspapers http://dream.lib.utah.edu/digital/unews/
Colorado's Historic Newspaper Collection http://www.coloradohistoricnewspapers.org/
Yale Daily News Historical Archive http://images.library.yale.edu/digitalcollections/YaleDailyNews.aspx

I then created samples of editions in JPEG and PDF formats to help determine which format best fit our collections needs. After examining the OCR of both file formats I found the JPEG files that had OCR done by ABBYY Fine Reader when imported into CONTENTdm to be more complete and accurate then that done by Acrobat when creating the PDF files. The display and navigation of both formats are very similar so we decided to use JPEG based on the OCR.

The metadata template was then developed by looking at the needs of the collection, standards and examples of the fields that other institutions had used for similar projects. A schema was then decided upon and followed for creation of the sample collection.