Showing posts with label D0DA12. Show all posts
Showing posts with label D0DA12. Show all posts

Monday, October 15, 2012

Late Breaking...

I just wanted to post two more entries in D0DA this year that I didn't see until after work on Friday:

I'm posting them a bit late, but Vive la Day of Digital Archives, right?!?!!

Friday, October 12, 2012

Cloud storage and digital archives


Hello everyone! Since I am new here, I better introduce myself, but briefly. My name is Sarah Kim. I am a PhD student at the School of Information, the University of Texas at Austin. For several years, I have been exploring people’s everyday digital record-keeping practices. Personal digital archiving as a form of long-term digital record-keeping and value-determination practice is the phenomenon that I particularly focus on in my research. I have been asking people how they live with their digital documents. Today, I would like to share one of the questions that I have been thinking of based on my (personal) digital archiving research: Cloud storage and digital archives.

During the interview, I asked participants what they would pick first to rescue if there were a fire at their home, besides living things. Many participants mentioned an external hard drive or a personal computer. Although their answers may be influenced by other record-keeping related interview questions, if someone asks me the same question, I would think of taking my external hard drive that functions as my own digital archives (not as mere back-up storage). Digital documents (including pictures) stored on that device are vital for me to rebuild and continue my life. 

Recently, many IT companies are offering cloud computing services as massive data storage. IT researchers and practitioners often call cloud computing a new paradigm for computing. In fact, many people are already using cloud services to conduct their work and/or non-work related activities (e.g., Gmail, Google Docs, Dropbox and many others). Cloud storage has a potential as a future platform for personal digital archives (as well as digital storage of memory institutions  — Interesting survey results of National Digital Stewardship Alliance member preservation storage systems: http://blogs.loc.gov/digitalpreservation/2012/01/partly-cloudy-trends-in-distributed-and-remote-preservation-storage-more-results-from-the-ndsa-storage-survey/).

This makes me curious about how my participants’ answers will change once they actively start using cloud storage to keep and preserve their personal digital documents and furthermore how our digital archiving practices will change with the new technology?

(Well-known New Yorker Cartoon by Mick Stevens, Published November 21, 2011)

It is highly likely that more people (and memory institutions) will be interested in using cloud storage as their digital archives considering benefits associated with it and the overall trend in IT industry. Cloud storage, (expected to be) maintained and monitored by IT experts, could be a relatively more secure place, considering the technical vulnerability of  more conventional digital storage media that many people are using such as external hard drive. Cloud storage services offer other useful functions such as sharing documents with others, synchronizing digital materials between different devices, tagging, and so forth. Also, cloud computing is still in early stages of development in general.

There are, however, many questions to ask to clear the cloud of cloud computing. For example, concerns for privacy (We are very familiar with horror stories about data hacking, personal information selling, identity theft and so forth) and feeling of losing control over their personal documents (Who owns what on the Web?), and building trust between users and services providers (How much can we trust the work ethics and long-term sustainability of these commercial services?) remain vital issues that we need to think of. 

From an archival perspective, I think, memory institutions’ (especially archives) working experiences with cloud storage service providers can offer a great insight into how we can inject archival thinking (e.g., what archives means, values of documents, and so forth) and practices in the design and development of these services. 

Thank you for reading.
Sarah Kim 
(Personal digital archives research blog: http://personaldigitalarchives.blogspot.com/)

The William Blake Archive on Day of Digital Archives

Make sure you read the posts over at The Cynic Sang: The (Un)Official Blog of the William Blake Archive:

Day of Digital Archives: Artist's Collections

Well, this year's Day of Digital Archives ha been much more successful for me than last year (I broke my elbow in a cycling accident that day and spent much of it loopy because of the pain pills. I did type out a one-handed blog post but I don't think it ended up being coherent.) This year I want to talk about an artist's collection we've been working on for a bit at UO.

The Tee A. Corinne papers are one of the many hybrid collections we have in Special Collections and University Archives at the University of Oregon. Tee Corinne was a lesbian visual artist, writer, and activist who explored female sexuality in her visual and written works. Upon her death in 2006 she left her entire estate, including the rights to her literary and artistic works, to the University of Oregon Libraries. Owning the rights is nice because, once we've done our initial processing and preservation work on the files, we don't have to worry about any rights issues when providing access to the digital objects.

However, before we can even start worrying about access to the materials we've had to devise a plan for working with the digital records. When UO received Tee's collection in 2006, it included a laptop and a desktop computer as well as removable media containing various works and papers. At that time, the UO did not have well-developed procedures of workflows in place for ingesting or otherwise processing digital objects. The files were pulled off Tee's computers and the various media and moved over to library servers, but nothing else happened to them for a number of years. In the meantime, there was a gap of more than a year between the time my predecessor (the first e-records archivist at UO) left and the time I was hired. The Tee Corinne e-records were left on the servers and until now I haven't been able to work with them at all.

When I started my initial assessment of the digital portion of Tee's papers, my first task was to try to gather all the digital objects from the collection into one place on the server. Because of the lack of workflows when the collection was taken in, the digital objects ended up in a number of different places on the server. Although I think I've managed to round up most of them now, I still run across stray files that have to be added in with the others. When we started the project this summer, we identified 65,328 digital files we knew came from Tee's computers or from the removable media in her collection. Although I would love to be able to declare that all those files in fact belong in Tee's collection, she shared her computers with Beverly Brown, her lover, whose collection the UO also owns. In addition, Bev Brown was the founder of and was heavily involved with the Jefferson Center, an organization whose records the UO holds as well. Once we started looking at the files from Tee's computers, we realized that her files, Bev's files, and files from the Jefferson Center were all mixed together. The organic file structure the women were using did not clearly distinguish among these three separate groups. Often a single directory will contain files from all three collections. This has slowed down our processing: we're trying to develop some content-based filters so we can do some batch sorting of the files. Most of the textual documents were created in version of WordPerfect, so we're also working on batch converting those files. In addition, of course, we're having to do a lot of renaming so that the file names of the preservation copies don't have any of the potential trip-ups you see in organically-named files.

The most interesting challenge in dealing with this collection, however, has been the photographs. Photography was one of the many media in which Tee worked, and she made extensive use of Photoshop. Sometimes she created prints of several digitally-altered versions of a single photograph; we are often able to match physical prints with digital files, but in some cases we have digital photographs for which no physical print exists or vice versa. Tee also tended to revise her photographic series depending on the context in which she was exhibiting or publishing them. This means we sometimes have several different series of a single image or group of images. The series may or may not be consistent; that is, sometimes a series of images was published in one form in on place and in a different form somewhere else. In the digital files, this means that in some cases we have many duplicate copies of a single image (if Tee organized the files based on the various publications) as well as multiple different versions of an image. We would prefer not to transfer multiple copies of a single image onto our preservation servers, but we do want to preserve the different versions of the images because we feel these are an important artistic statement. Sorting out the files themselves has proved to be an enormous challenge, however. Luckily I have a team of graduate students and volunteers who are working hard on this (as well as other) projects.

What have I learned from my work with this collection so far? Obviously, documentation is a hugely important factor when you're talking about a born-digital collection. One of my main problems right now is the lack of documentation from previous work that occurred with this collection (however cursory that work might have been). I'm trying to document every step I take with these records so that my successors have a clear picture of what has and hasn't been done with the materials. It's also important for the digital archivist to be involved in the donation process if at all possible; this helps lessen the amount of triage work you have to do when the born-digital records arrive on your doorstep.

Day 29



Today was my 29th day of work as the first Records Management Archivist at Johns Hopkins University.  My job encompasses two areas that overlap frequently but not perfectly: management of university records and management of born-digital archival materials, regardless of whether they originate within the university or with external donors. This combination of roles is relatively common in our profession, but I haven’t personally experienced it long enough to evaluate it critically; perhaps that will be a topic for next year’s Day of Digital Archives post.

My first 6 weeks have coincided with the processes of annual reviews and setting individual goals in our library. Although I was initially wary of having to set annual goals so early in my tenure, the timing has been fortuitous because I have been planning my activities for the next 12 months – which I would be doing at this point in a new job anyhow – at a time when my colleagues are all thinking similarly. And if there’s one thing that I can say with certainty about the next year, it’s that it will involve a lot of collaboration: with other archivists, with curators, with developers, with metadata specialists and with project managers, just to name a few.

For the next few months, I will be assessing the current state of our institutional climate, our capacities and our collections as they relate to acquiring, preserving and providing access to born-digital archival materials.  Next, I will be working with my colleagues to determine what capacities we want to develop as an organization. Do we want to do forensic captures of media-based accessions? What kind of preservation activities do we want to undertake? What types of functionalities do we want to build into our digital repository? How do we want to provide access to our materials? Although there will be many details left unanswered at this stage, I hope to be able to address these and similar questions at a very high level within the next six months.

Finally, I will spend the rest of the year developing a three-year road map for how we can get from where we are to where we want to be – or at least, from where we are to moving purposefully and surely toward where we want to be. This will involve identifying gaps in our current technological and human capital, and proposing ways to bridge them.

Of course, the day-to-day activities of the archives will not stop for a year while I figure all this out. Prior to my arrival, no one in our department was charged with focusing to this degree on all the issues surrounding born-digital materials. However, like many institutions, we had still been acquiring them for some time. So while I am doing high-level analysis and planning, I will also be carrying out the day-to-day activities of accessioning and caring for our materials as best I can with the resources currently available. 

I have already made a few changes that bring our activities more in line with, for example, the minimal levels of digital preservation outlined in a recent proposal from NSDA.  Specifically, I have instituted the use of LOC’s Bagger tool to generate file manifests and fixity information according to the Bagit specification at the time of acquisition, and I am working with library systems to transfer our current holdings to a storage space where they can be more appropriately managed. 

However, I don’t anticipate many other changes in our procedures in the next year. This means that for the next 12 months, we will undoubtedly continue to do some things in ways that I know could be improved.  However, when I do begin to make radical changes in our procedures, they will be guided both by best practices and by our own organizational needs and goals.

New perspectives can translate to new opportunities

I’m the Digital Collections Archivist at Kennesaw State University in Kennesaw, Georgia. Kennesaw State is the third largest public university in the University System of Georgia, with a current enrollment of 23,103 for the Spring 2012 semester. Founded in 2004, the Archives consists of one full-time Archivist (me), a part-time Archivist, who also works half-time in the Bentley Rare Book Gallery, and an Associate Director. From 2004 to 2008, when I was hired, the Archives was staffed solely by the Associate Director. For the 2011 Day of Digital Archives, I created a photo essay to illustrate the different roles and responsibilities in my position. It was appropriate for the time, because our department was growing and expanding. We merged with the art and history museums on campus to form a super-department: Museums, Archives & Rare Books. This year, though, feels like one of retrenchment. We lost a long-time member of staff at the beginning of the year. The redistribution of her workload among the remaining staff brought fresh eyes and energy to some long-standing issues. We were able to use it as an opportunity to make significant progress on projects that had been stalled. In light of the difficult economy and constant budget cuts, I think similar organizations will find our actions of interest.

Writing it down 

At the beginning of the year, we hired a Records Manager, the first in the history of the university. She’s been working to understand and document the workflow of records creation and disposition across departments in the university. We don’t have an enterprise document management solution, so it’s been quite an undertaking. As part of her duties, the Records Manager has also inventoried records at our off-site storage vendor, identifying and transferring materials with historical significance to the Archives. Although the mission of the Archives is to collect and maintain university records that document its activities and history, we found that we had no recognized authority to transfer records without the consent of the department or division head. Trying to find someone who was willing to accept responsibility or to grant permission was an exercise in futility. After trying to track down one division head over the summer, it was decided that we needed to seek the authority to transfer records deemed to fall within our collecting policies. The problem was we had few written policies.

The department formed a policy committee with representatives from the museums, Rare Book Gallery, and the Archives to develop unified collection management policy. We found that we were able to use the same language and concepts, adding specific examples or language for situations unique to each unit. After several iterations, the committee was able to create a collections management policy over the summer. It’s currently awaiting final approval by the Chief Information Officer before being implemented. Once this is in place, we are ready to submit the transfer authorization proposal to the President’s Committee for approval. The completion of the collection management policy spurred interest and development in additional policies and procedures, including reproduction, access and use, and registration, as well as related forms. We’re currently working on creating copyright policies. My particular focus is developing guidelines to help users to understand copyright restrictions and to make responsible reproduction decisions.

Clearing it out

Looking at ongoing problems with storage space, both physical and digital, we created an ad hoc committee to review materials and make decisions regarding disposition. One of the first problems that we identified was a large amount of supplies and resources that had been amassed “just in case.” These included outdated or broken equipment, unnecessary or unusable supplies, and donations that did not meet our collecting areas or interests. The process of clearing out the space allowed us to reorganize supplies, to order new equipment, and to relish the sense of accomplishment. We used this momentum to tackle the shared drives and digital repository, both of which had become dumping grounds. The same committee developed policies to govern the shared drive, as well as file naming conventions. Using these new documents, we began a clean-up of the shared drive, which amounted to removing approximately 60 MB of duplicate or unnecessary files. The process turned out to be so easy that the committee offered the services to one of the museums. We were also able to incorporate elements into an outreach session on email best practices and plan to offer the service to other departments on campus.

Building it up 

The university is coming up on its 50th anniversary in 2013 and the Archives was approached to digitize historic images and to make them available for users. Currently, we rely on Archon to provide public access to our records and small files, such as oral history transcripts and low-resolution images. It was decided that it is inadequate to provide access to the high-resolution images required for the anniversary. I was tasked with comparing systems and making a recommendation. After much research, I determined that DSpace would best meet our requirements. We’re currently working with campus IT to implement a DSpace instance. As part of the DSpace project, I mapped the workflows of the Archives and identified current and future technology needs. This plan can now be used to ensure that we make strategic decisions based on demonstrated needs.

In addition to implementing new systems, we’re also focused on improving our current products and services. Archon was originally populated by importing data from our old CMS. It contains many records with minimum information. As part of the general commitment to bring consistency to our records, I’ve initiated a project to enhance the catalog on a record-by-record basis. This also allows me to check new accessions and add them to existing collections when appropriate, as well as to verify location and beef up the MARC record in the library’s OPAC. The project has already revealed some MARC mistakes and location errors.

By mapping the Archives’ core functions and relating them to technology needs, we are able to offer products and services of higher quality and with greater efficiency. We were also able to use our clean-up as a template for new services. While retrenchment may not seem as exciting as rapid expansion, it can still be an opportunity for growth and improvement. Please feel free to contact me if you'd like to ask any questions or follow up. You can reach me at agraha31 (at) kennesaw (dot) edu.

Many thanks to Gretchen for providing the opportunity and forum!

More Blogiverse

Here are some good reads that weren't necessarily created for Day of Digital Archives, but are related just the same:

I'll be heading home soon, but our west coast and other international friends can keep the conversation going!

Across the Blogiverse, part two

There are lots of fantastic posts going up all over the place to celebrate Day of Digital Archives. If you haven't yet, be sure to check out
  Happy Reading Everyone!
Instead of writing a traditional post for Day of Digital Archives this year, I'd like to link to the slides for a presentation I just did earlier this week for my library. Earlier this summer at UVa we went through a dramatic series of events related to the resignation and subsequent reinstatement of our University President. As Digital Archivist, I was involved in a larger library effort to create an archive of materials related to events. The campus community became quite active and vocal and organized several meetings and demonstrations using Twitter and Facebook -- a mini version of the so-called "Twitter Revolutions." My role, as I saw it was to act fast and try to save as much evidence of these activities as I could. The process was a challenge for many reasons, some technological, some legal, and some just human (I do need to sleep at some point, you guys!).

The slides themselves are embedded below, but if you choose to view them at SlideShare you can also view the slide notes, which are basically the narrative of what I said. It's much more informative with those!


Across the Blogiverse...

Welcome to Day of Digital Archives 2012! If you haven't already, be sure to follow the #DayofDigArc hashtag all day.

In addition, our colleagues have been blogging away.
 Be sure to check in over their for insights on their work and really great digital archives!

The StoryCorps Archive: A Brief Introduction

We're happy to take part in the Day of Digital Archives! For more information about StoryCorps, please visit storycorps.org.

Many people know StoryCorps through listening to our broadcast pieces: brief but hard-hitting two or three-minute nuggets that tell compelling stories. They may be tear-jerking or funny; some relate the story of an overlooked historical moment and some simply convey interesting anecdotes or depict unique characters. However, not all of our listeners realize that each of these clips is edited from a 40-minute long interview (which typically takes the form of a conversation between two people who know each other, like a pair of friends or a father and his daughter), or that, about nine years since StoryCorps’ founding, we have collected some 45,000 interviews—that’s 30,000 hours of tape, all in CD quality WAV format (44.1k, 16b). Transcribing all this audio would require a daunting amount of resources, and, consequently, the vast majority of our interviews lack transcripts.
 

Despite this issue, our cataloging practices do provide multiple points of access—including keywords. We partner with the American Folklife Center at the Library of Congress, which guarantees the long-term preservation of our collection and its associated data. With their help and using a draft form of the Ethnographic Thesaurus, we have come up with a customized list of terms that we use to tag interviews. These keywords allow for the possibility of non-linear, subject-based searches within the Archive, a process that creates connections between interviews with seemingly little in common.

Since these interviews represent a vernacular history, a record of events ranging from world wars to family holidays that is uniquely rich in affective detail, it seems important to enable their ongoing accessibility to researchers and to the public. But what, exactly, do we do with this massive amount of digital audio? We don’t yet have a reliable audio search tool that we can use to delve into each file’s contents.


To address this overarching question, we’ve created a few tools for internal users to make the content more digestible. The facilitators who are present during the recording of each interview take handwritten log notes detailing each interview’s content, which we scan and include in each individual record. As they create database entries for interviews, the facilitators also transcribe five points from their log notes, which users can search in our database system (a customized Drupal-based database). Users can also scan through the interview’s full audio file using these time-coded notes.
 


Our content searching has enabled us to expand access to previously unheard moments in the StoryCorps Archive. As part of a collaboration with radio producer Krissy Clark, we combed the collection for interviews that mentioned specific locations within the diverse neighborhoods of Lower Manhattan. Krissy edited full interviews into over thirty short excerpts and created a geotagged sound ramble through downtown that we presented to attendees of the New Museum’s Festival of Ideas in May 2011.

We’ve also been able to establish two partnerships with linguistics researchers—one team at MIT and one at Oregon Health Sciences University— that represent a model of collaboration that we hope to pursue further. Our partners at MIT’s Lincoln Laboratory study African-American Vernacular English. As part of their project, researchers took a representative sample of StoryCorps interviews and subjected them to computer analysis, specifically focusing on speech and dialect patterns. As a result of their research, they generated transcripts that we were able to add to our own records for future researchers to use.

Recently, we hosted an advisory summit supported by the Alfred P. Sloan Foundation entitled “Reimagining the Archive.” Data scientists, statisticians, archivists and librarians, oral historians, and linguists came to StoryCorps’ offices for a day of brainstorming the possibilities inherent in a large collection of digital audio and its associated metadata. Panelists encouraged us to sponsor a “hack day” that would allow innovative techies to forge new paths through our data, to build out widgets so partner organizations could “curate” their own subsets of the StoryCorps Archive, and even suggested that we create a mirrored server where large institutions could run processes on the entire archive to analyze speech patterns, word choices, or metadata. We look forward to piloting some of these ideas and to working with new partners in 2013 – hopefully we’ll be able to show some results on next year’s Day of Digital Archives blog!

Monday, August 13, 2012

D0DA 2012

The Day of Digital Archives is making a return: October 12, 2012 you can help raise awareness of digital archives and digital preservation issues through this community blog/twitter project. On this day, archivists, digital humanists, programmers, or anyone else creating, using, or managing digital archives are asked to devote some of their social media output (i.e. tweets, blog posts, youtube videos, etc.) to describing their work with digital archives. By collectively documenting what we do, we will be answering questions like: What are digital archives? Who uses them? How are they created and managed? Why are they important?

Last year's Day of Digital Archives featured more than 50 bloggers (many of whose posts can be viewed here through the archives) and more than 700 tweets. The topics of posts and tweets covered a broad range of activities from early discussions of the need for particular tools to announcements of completed products. Others used the platform to discuss things like education and training, collaborative initiatives, gaps in tools or shared knowledge, or the activities involved in planning or carrying out projects.

Do you create, manage, or use digital archives? Would you like to participate? Well then, drop me a line at gretchen[.]gueguen[@]gmail[.]com with your contact info and I’ll add you to the list! You could contribute in a couple of different ways:

1. Create a blog post at http://dayofdigitalarchives.blogspot.com/ for the 12th of October about some aspect of your work with Digital Archives on that day. It could be a really specific exploration of a single activity on that day. Or it could be a broader topic not really related to that specific day (What kinds of tools you could really use to process a born-digital collection).

2. Write a post to your own blog similar to that described above and post a trackback to the Day of Digital Archives blog or let me know and I'll post a link to it.

3. Tweet throughout the day about your work with digital archives using the #DayofDigArc hashtag. Hope to hear from you soon!