Showing posts with label metadata. Show all posts
Showing posts with label metadata. Show all posts

Tuesday, July 28, 2020

From Cancellations to Coding: Pandemic-Centered Tech Topics on Day Two of the OBS/TS Summit 2020

So far, day two of the summit has delivered fantastic programming. I wish I could attend it all! The final virtual event takes place at 6 PM EST tonight. This morning my two favorite sessions both dealt with the new realities we are living in post COVID-19 closures, touching on this from the perspective of budget cuts to work from home workarounds. Here were my takeaways:

Top Left to Bottom Right: Gilda Chiu-Ousland, Wendy Moore, Heather Buckwalkter, Anne Lawless-Collins.

TS Resource Management Roundtable: Budget Cuts & Collecting Pivots
I was really on the fence about which of the earliest morning sessions to attend, and I am so glad I selected this one on resource management and collecting pivots. Wendy Moore from the University of Georgia Law Library led the discussion with a powerful statement that really summarizes the entire roundtable and the timeliness of the topics:
"Crisis can lead to LOTS of creativity."
What followed were introductions from each of the panelists including Heather Buckwalter, Gilda Chiu-Ousland, and Anna Lawless-Collins. Each shared the state of things at their institution, the fallout from COVID-19 closures including the stopping of shipments and the addition of online study aids and other e-resources to help students and faculty get through a quick pivot to virtual learning, and the budget (if they had %'s or figures yet) that they are each facing for fiscal year 2021 and 2022. This session (as with several from day one of the summit) was not recorded to allow attendees to feel more comfortable sharing the details and situations of their library, law school, or larger institution. Two polls were executed in the larger Zoom room before dividing into smaller groups for more personalized and in depth discussions. The polls were very interesting, revealing many of us still do not know our budget, or have vague %'s that are yet to be approved, and that the majority of us are cutting print journals more than any other area of our collections.


In the smaller groups, attendees were better able to share their own situations, including some very creative strategies for how to negotiate with vendors, what data they are using to make those decisions about what and how to cut items from the collection, and what they have already or are planning to cancel to meet the demands of the coming fiscal year. There was a big focus on mitigating expectations of faculty and other stakeholders, and many were open about having these difficult conversations with their faculty members related to monograph acquisitions and with their institutions related to print course reserve materials. Overall an excellent program that was really open to sharing their situations so we can all learn from one another and continue best serving our library users.

Hot Topic: Technologies We Use

Presented by Jesse Lambertson, this session was more of an open discussion than a straight-forward presentation. Sharing his own library system as the beginning example, Lambertson pitched questions to the audience with lively responses in real time and invited members to un-mute and speak to their specific system challenges in the work from home environment. It was really interesting to hear individuals sharing the pros and cons of their various integrated library system platforms once they were catapulted into teleworking. The clear up-side to having a web-based interface was the ease that these librarians and their staff could quickly pivot to working from home without the hassle of using VPN or requiring remote desktop. These included those using TIND and Alma to name a couple. Several of us still working with iii's Sierra were able to join in chorus about our struggles in working from home with spotty VPN support and the differences in Sierra web as compared to the desktop client.
Presenter Jesse Lambertson screen shares Python script snippets hack for working with CSV data.
For importing and exporting records, both individually or in batches, many hacks were shared including creative ways use Marc Edit when working from home and the potential for more API's between Marc Edit and the ILS. It is of course that time of year when we are all gathering statistics. With much overlap from the previous session I attended, many of us commented we are accessing collection and user data much more right now to better inform decision making in a time of budget cuts. As a result, further roadblocks and workflow workarounds were discussed for various systems. Several attendees shared how they query their system for cataloging and other statistics, the issues they experience in the format of the data they pull out, and the obstacles that come with trying to do this type of work from home or with very limited access to the library. Many individuals (myself included!) are periodically retrieving data from their systems, exporting it at txt or csv files, and then taking it home on laptops of flash drives to be able to spend more time with it when teleworking. However, and few shared more innovative approaches to both massaging data as well as collecting and sharing it. Lambertson shared a highly creative approach using Python scripts to automate certain aspects of the csv to Excel conversion of his data. Another attendee shared their library's customized Google Sheets dashboard which pulls data from the ILS into the same location as reference transactions statistics (populated by Google Form responses). A truly fantastic session with lots of open dialogue between attendees. I am so glad I attended and I can't wait to see and hear how the experiential system and data approaches our members are working with now unfold in the coming months and years as access to our offices and systems remains largely unknown during a pandemic.

Tuesday, April 28, 2020

Power Projects for Quarantined Librarians


Some of us are approaching the two month mark of our library's closure to the public. Though it has definitely had ups and downs, I have found it has helped me better carve out time for professional development activities, dedicate more of my day to clean-up projects I have just not had time to follow through on, and even start new projects with the help of colleagues who found themselves needing a little more to do from home. On the heels of Travis' post about transitioning technical services staff to working from home, I hope this post will elaborate on that topic to give specific project examples for those working in collection services, technical services, or metadata and archives related positions.

ILS Database Maintenance
Projects of this nature could range from starting, continuing or finishing the cleanup of small or large sets of records.
  • Complete outstanding cleanup projects a for smaller sets of records. For me, a cleanup I started last fall (LL.M. Theses Collection) which I previously blogged about was the first on the chopping block.
  • Facilitate the cleanup of collections, like course reserves, for other departments. Our role was simply to assist by creating lists, then formatting those in Excel. Access services staff can use that list to check for faculty members no longer with our institution, and then remove the instructor and those related reserve item records from the ILS.
  • Learn more about your system and the tools you can use to better care for it. There is no shortage of webinars right now. In addition to system specific sessions, there are many free sessions focusing on record management or data tools. Some of my favorites have been Terry Reese's "MarcEdit Shelter-In-Place Webinars". The latest one, number 6, is on regular expressions.
IR Micro & Macro Clean-up
Repositories like our own in Digital Commons have issues similar to those inherent to long-standing ILS's: without regular care and feeding, the structure becomes less organized and the data for records less consistent.
  • Review your repository site map and make big picture structural adjustments. Our IR has been around since 2006, and many of our earlier events (like conferences and symposium) were added before we adopted using the event types. As a result, all "events" before a certain date were actually article-type records, and all of the "events" after that certain date had completely different series structures as well as different metadata fields. To resolve this some time was spent almost manually moving the articles over, one by one, to give them the appropriate fields in their new home. It has taken lots of effort, but in the end it will make all the difference in the discoverability of events of the same series. For more info on this topic, see a previous blog post about IR Metadata.
  • Harvest digital media to expand your individual item content. Similar to the LL.M. Theses collection cleanup mentioned above, last fall another colleague and I began a collaborative project to use scripts to capture all of the metadata and publicly hosted digital-born image files from our law school website into spreadsheets of data. Using these spreadsheets after a little minor formatting of the cells I can now batch load the images pretty quickly. Though this work is still in progress, it has been a perfect project for both my colleague and I to tele-work on since the repository and the website are each accessible from home.
  • Learn more about your system. The BrightTalk Digital Commons sessions have been wonderful to watch both live and recorded versions of (learning more about native streaming has really come in handy!), and many of the past CALICon sessions which are all freely available as videos online have also been great sources of learning how others are using repositories and what else we can do with our own IR.

Cataloging Collections & Archives 
Be it physical items, special collections, or virtual equivalents, the building being closed has not stopped the number of items that need our attention. Even without new items, existing items can always use accessibility makeovers!
  • Archive virtual events. As we all know, although some events have been cancelled completely due to closures, most have opted to go virtual in one way or another. At our law school, faculty colloquium have continued occurring in Zoom, and even a conference was hosted entirely online (with more attendance than our physical space could have accommodated). I have continued collecting materials for archiving these events both for when I return to add to the physical special collections, and to add them as I normally would to our repository "conferences" series. In the absence of printed programs, I have saved PDF "prints" of email programs, and instead of photographs of the room or panelists in real life, I have taken screen captures of Zoom rooms at the highest quality my computer will allow.
      Thumbs of 15 photos of book spines for a colleague.
  • Catalog items from home. I did not have too many outstanding items to catalog when I left the office to set up my home workspace. The items that I did have, I brought with me in a small box. Most ILS have a web-browser accessible entry point, and although I cannot complete all of my tasks from home (like data exchange) I can still catalog! I was able to catch up on a few items this way, and honestly spend more time doing detailed original cataloging that I might have rushed through in the office. As one of the essential employees (those physical backup tapes don't change themselves!) I have been able to grab a couple more items as needed on my brief but weekly run into the office. For those colleagues not coming into the office at all, I've taken photographs of items to share with them so they can continue their work from home, even without the items in hand.
  • Make items more accessible. This could take lots of forms. One project we are excited is finally underway is the OCR-ing of hundreds of already digitized documents. The PDFs were not text-searchable, but thanks to one staff member and a couple of librarians we have created a very effective workflow and are making great progress to provide more accessibility and in turn discoverability to archival collections like student directories, law school magazines, and historical strategic plans. Transcription is another option if you have more audio or video content. It can be tedious but the effort goes a long way to making items available to a wider audience online. Marketing your collections and archives is another way to share them with the world. Blog for your library about physical items to help patrons or the public feel more at home, even from a distance. In doing so I've used it as an opportunity to get to know our archives better. Advertise your digitally available collections through organizations at the state, regional or national level. Everyone is searching for free educational content right now, so share what you have to offer. Many org's have made calls for this type of content to spotlight, and others will probably thank you for sending ideas their way. 
Using Adobe Acrobat Pro on Law School issued laptops, collection services staff could batch enhance scans and perform optical character recognition (OCR) to make PDFs text-searchable.
 
Professional Growth & Contributions
If you are still finding yourself, your colleagues or supervisee's lacking things to do, sign up for a course, attend a webcast, document your projects and turn them into articles or presentations.
  • There is no end to organizations sharing webinars online right now. Even if you cannot fit all of the live events into your teleworking schedule (for some reason many of them have taken place at the same time, and in the same platforms!) you can still register for them to receive access to recording links later. In addition to the links I shared above, I've also really enjoyed the two courses I took over the last two weeks from Midwest Collaborative for Library Services. It never hurts to refresh your memory of certain topics that you might not have had dedicated attention to give since library school, or to learn something brand new! 
  • Document everything. Each time I begin to make progress on a new special project, or fine tune a workflow at my institution, I approach documentation as if I were going to present it later as a conference session, workshop, or article. With many conferences announcing they are going virtual, it has never been a better time to submit proposals without the hesitation of travel logistics. If you have never published before, now may be the time to share how you did that certain something with a journal or in your organization or SIS newsletter. The opportunities are limitless, so keep that in mind no matter what projects you are undertaking. 

Tuesday, January 14, 2020

Cleaning Up Messy Records: Uncovering Match-Points in ILS and Repository Data

Many of us have our hands in multiple pots. Sometimes we are working with repository records, and other times records in our ILS. While embarking on what I mistakenly thought would be a simple series of tasks (linking from 856 fields to our freely accessible, digitized versions of the same items to our IR), I happened to uncover much messier data than expected. What a perfect opportunity to do some house-cleaning! In case you have undertaken similar work, and are considering comparing, cleaning, and (eventually) updating records in multiple locations, here are a few tips and resources I have found helpful on my own journey:

Have good list from the repository. My initial cleanup of Digital Commons items was fairly straight forward. I sat down with our International Law Librarian Anne Burnett to talk about the collection. We batch-downloaded the series as a spreadsheet, and updated the fields that did not match so that they did (example: type was not the same for each, some were articles, some dissertations - this one was easy, updated them all to be "dissertation" type). Then I saved it, and re-uploaded the batch sheet.
Have a good list from the ILS. The same set of items in our library catalog was more difficult to get a solid list of. We are in Innovative's Sierra, so I needed to use "create lists". This time sitting with Associate Director for Collection Services Wendy Moore and I generated a list of items (not bibs). We made several lists, because we quickly discovered that what we thought our control field was (the 502 note) was inconsistent. Catalogers had changed over the many years these records were created, and at a certain point the 502 had changed from having only two periods in this item's abbreviation to having three (ex. LL.M. to L.L.M.). This minor flaw made things a bit more difficult. We ended up instead using the donor note field to get closer to the ideal number of items from my repository list. In the future the location might also be a field to use for pulling this sort of list - but part of the ILS record cleanup was updating location codes, since items had recently been shifted from reserve to the basement. 
Fix your controls so they actually work. I ended up going with the list that had the highest item count (although now I had more than what was in my Digital Commons list) and updating these
records first. Now they will have consistent 502's, and correct item locations. To verify the proper 502, Wendy and I consulted our office copy of the AACR (Anglo-American Cataloging Rules) -
  • "Section 2.7 B13 Dissertations". Although LL.M. was not listed as a specific example, the rule states that you use: "Thesis followed by a brief statement of the degree for which the author was a candidate (e.g.. M.A. or Ph.D.), the name of the institution..., and the year in which the degree was granted." 
To do this I used a combination of Global Update in Sierra and manual/individual edits. I added 856 fields to the records as well with subfield u (linking to the Digital Commons series landing page for this collection) and subfield 3 (for the text I wanted to appear as the hyperlink access point). You could do some of this in Rapid Update as well, if you're feeling extra confident! - Note: In Sierra I had to start with a list of item records, but at a certain point I had to re-run my partially cleaned-up list to create a new list of bib records. I could only globally update fields in the bib record with a list of bibs (a list of items would not work!). 

Export a better list from the ILS. Now that I had a proper list from Sierra, I exported it to a text delimited file, then imported it into an Excel spreadsheet. I was ready to compare it to my repository spreadsheet and figure out what what missing in Digital Commons.To figure this out, I started by doing a "save as" of each spreadsheet, then narrowing both sets of data down to only the fields that I could use to compare between the two: Title, Author, and Publication Date (year was really what I was looking for). This presented further problems - not all titles in Digital Commons looked like what should be matching titles from the ILS records (ex. many repository title fields had been entered in all caps!). For Author, Digital Commons separated last and first name fields, but in the MARC records this was a single field. For Publication, the formatted date in Digital Commons records was very detailed and specific, while the only match-point in the ILS records was the 260 field (included publication date at the end as the subfield $c) - major thanks to our Collection Services Manager David Rutland for knowing this one off the top of his head! The 502 might again prove useful if the 260's were too difficult (since the "year in which the degree was granted" appeared at the end of this field for each item).


MarcEdit, OpenRefine, and beyond. At this point in my process of this particular project I am playing around with a combination of editing in MarcEdit, as well as a number of tricks I am reading up on for OpenRefine cleanup. If you are totally new to OpenRefine, there is a really handy wiki with screencast intros that I appreciated before diving in. So far, an extremely helpful resource I have come across has been Comparing Two Sets of Data in OpenRefine How-To. This entry shares step by step how to "Normalise titles to do comparison" using three main transformations. For my particular set of data, the value.fingerprint transform has given me good results, removing case from both sets of titles:

There is also an excellent page with more information specific to working with and cleaning up dates. I am still working with the cleanup of this set of items, but even though it is a work in progress this has been a wonderful learning experience. Each time I work on it I learn something new! I am excited about the things I have figured out in this process that can be applied to other sets of items in both our repository and library catalog records in the future. I'd like to thank several of my colleagues at UGA Law Library for providing various pieces of this project's puzzle. Without them I would not have made it this far with these particular data sets. Thank you Anne, David, and Wendy for all your context, tips, tricks, and sharing your experiences with this collection.

What types of cleanup are you doing with your library's data? What tips and resources have worked well for you? Please share with us in the comments! 

Monday, November 25, 2019

Here Come the Bots: Six Tips When Designing Your IR's Metadata for Improved Discoverability


Last week I attended a webinar about "the science of discoverability". Although it was aimed at librarians working with institutional repository (IR) content, it was an excellent reminder that the many best practices I followed as a web developer for our law school's Drupal site were applicable not only with repositories but also with LibGuides (and any other pages we wanted Google to find). Here are six tips to deploy when designing metadata for the bots and increasing your site's discovery:

  1. Title Fields Are Important! In fact they are perhaps the most important field of any object or event metadata in your repository. Working as a web developer this was something we struggled with when other users would create webpages. The title did not always match or identify the content. Later on they inevitably call or email to ask why it isn't showing in Google's search results when they put in keywords that they think and assume will definitely retrieve their exact webpage that was literally just created - of course it doesn't work. Almost always the keywords they wanted Google to identify were not in their page title field (or URL). The same rings true for IR content. No matter how many other fields have the data or keywords, if the title doesn't it probably isn't good enough to be retrieved by Google (unless you have big bucks of course - then you can use Adwords to pay your way to the top of that results list as a sponsored item...but I doubt any of us have that kind of money for SEO, hah!).

  2. HUMAN-Readable Is Better. This is not your library catalog. Your ILS is a (mostly) closed-off system. It was engineered by ONLY librarians who have strict cataloging rules passed down over decades of meticulous fine-tuning with a field for literally every-single-possible-bit of data. IR's are not an ILS. In the same way Google is not your OPAC. They do not and will never function the same way. Sure, you can use some of the same operators, and you may even form similar strings in each of the search bars. The difference is that Google's algorithm is not a 100% known entity. Most of Google's users are performing natural language searches. Your I.T. or metadata librarian's cannot get into Google's back-end and tell it what you want, what fields to provide searches for, what weight to give certain types of results, or how to display your results list. Google's algorithm not only likes but craves HUMAN-readable, NOT machine-readable. Craft the content in your fields for any given item, event, or landing page with this in mind. You really should design the data carefully. And the key here is not to overdo it! 

  3. Don't Use Too Many Keywords. This relates to the last sentence of the last tip - don't overdo it. In addition to not getting overly wordy or technical in your fields, the field to especially watch out for is keywords. In Digital Commons there is a nice keyword field. When I first started adding content to our repository I no doubt went overboard with more keywords than I should have. Although too few could hinder discoverability, if the keywords are on point and you have two to four of them that are appropriate you will hit a sweet spot with Google's crawl. But beware of using too many. Google and other search engines will actually ping or potentially ignore your content (and in some cases as the webinar warned your entire site) for using too many keywords. Excessive metadata makes it assume this content isn't valid. So just be careful here. This doesn't mean you should never use more than four keywords. There may be occasions when less just won't cut it. Perhaps that one article or conference you just loaded is particularly interdisciplinary and really needs more terms. Keeping the majority of your content with three keywords or less will get search engines to take you more seriously and those few instances where you decided to use more keywords won't throw up red flags like twelve keywords for every single items in your repository would. 

  4. Frequency, Consistency & Longevity. I can't count how often I was asked as a web developer when Google would crawl our site. This is a mystery to most everyone, and while you can request through some of Google's Webmaster Tools for a re-crawl there is no guarantee the speed at which that will happen. One thing is for sure, you will be re-crawled more often the more frequently and consistently you update any site, no matter what site it is. Long periods of no activity may result in flagging you as a dead site so regular adding or refreshing of content is the key here. Another related factor is longevity. This is simply the idea that the longer a site exists the more time it has had to be crawled, to appear in search results, and as a result to increase site traffic. Then the cycle returns to the beginning since the more site visits you receive from organic Google searches the more your site should rise in the results list as your site and its content becomes more closely associated with a variety of searches over time. Obviously a brand new site will take time to get there, but after many repeats of this cycle (with the help of your frequent and consistent care and feeding) this will happen naturally. 

  5. Bots Like Quick Load Times. So since we don't really know when Google or other search engine bots will pay us a visit, how can we make sure that when they do they are finding us at our best? Load times are one big indicator. I know, I know... but there are SO many cool and flashy things we could embed into our content, right? Is that snazzy High-Res image of the latest guest lecturer too much for Google? What about our Issuu flipbooks of scanned symposia programs, or the YouTube video of the three hour panel? Each bit of multimedia needs a different approach here. If your IR system has native streaming this will help cut down on additionally embedded load times. If not, you may need to choose what is more important - the load time or the media keeping your traffic on your site. If traffic isn't a major factor, load times will increase by hyperlinking to the media instead of placing it on the page itself. The same could be true for embedded flip-book style PDFs. For images, as long as you use best practices for the proper resolution on the web you should not have to choose between a crisp, quality image and fast load times. Use the right format for image and other media files (choose MP3's for online streaming instead of WAVs of AIFFs). If you want or need to offer the highest quality original files to site visitors, hyperlink to that file's location instead of providing at their point of entry. This will keep load times up and still give visitors the option of access and retrieval. In the end, the faster your content loads, the more quickly it can be indexed. Bots are impatient - they are bots! Make them wait too long and they just keep moving. 

  6. Site Maps Are Critical, Especially for "Dead" Collections. So your content is now in tip-top shape! It has excellent human-readable title fields and abstracts. It has good keywords, but not too many of them. You've even managed to build a beautiful page of content enhanced with multi-media, but you've been careful to follow best practices for these files and your load time is great. Now there is just one problem - this collection is an archive! It just so happens as a librarian you have created a collection of items that will never grow again because it is historical. How can you possibly be frequent and consistent with this set of data? Will Google eventually forget about you (even if the collection exists over a long period of time) because there is nothing to update? No! Not necessarily - this is where your site's skeleton, the trusty site map, comes into play. Depending on the system you are using a site map may be generated for you as you create new content. It never hurts to revisit this though. Particularly for sites that have been around over a long period of time, the site map (generated for you or created by someone else) may be pulling titles and other structural and organizational information that is either no longer accurate or appropriate, or perhaps it is just not as good as it should be. Revisit your site map every so often as a regular maintenance task. It is essentially an outline of your site and all that it contains, and as such can indicate where a collection or series title is not descriptive enough, is too descriptive, or is just not human-readable. Think back to tip #1 and #2 for human-readable fields (especially titles). Page summaries can help here as well. When you conduct a Google search, if a result appeared but had no description at all for the page are you going to take your chances with clicking through to that result, or are you more likely to choose the result that tells you what you will find there? Make titles, related page summaries for what it is about, and if possible even URL strings make sense and describe what you will find there. Adjust your sitemap and related descriptive data as needed, and monitor how your site (hopefully) rises in results over time, as well as how your traffic (hopefully) increases over time. 
Have more tips to share with TechScans readers that were not touched on here? What has worked for improving your website or repository's metadata, and how do you optimize your content for search engines? Share with us in the comments! 

Monday, November 4, 2019

GLA Conference Review: Workshop on Digitization for Small Institutions


A while back I did a post called What About Conferences? aimed at newer members as part of our "Quick Question" series. In that post I specifically talked about memories and experiences from state and regional librarian organization annual meetings. One of those (perhaps the organization I am most fond of!) is the Georgia Library Association. I have been actively participating in this state association affiliated with ALA, ACRL, and SELA since I attended my first GLC (Georgia Libraries Conference) in the fall of 2014. It has always been a welcoming and lively group with a crazy awesome mixture of library types and individuals.

Co-presenting at GLC 2019 with colleagues
Szilvia Somodi and Marie Mize.
What I love perhaps most of all about GLC is that you will find all levels of librarians there (not just "faculty-level" with "Librarian" in their title). The very best sessions I have attended often come from library staff. As a librarian who worked a few public libraries while studying librarianship, and as one with past positions which until recently were entirely "paraprofessional" or I.T. titles this is where the on-the-ground, behind-the-scenes knowledge and skills are found and shared: at the local events candidly. GLC isn't pretentious or intimidating like some conferences and their crowds can feel. It is also not overly techy like many I.T. and web developer conferences I have experienced where all you hear is jargon that feels distant and mysterious. There is a beautiful happy medium at GLC where you can network, actually learn, and find encouragement to follow your interests and grow as a librarian without judgement. It is here that my love of libraries grew stronger, although it took me a few years to get comfortable enough to sign up for one of the pre-conference workshops.

Table of recommended project management
software from DLF's awesome wiki.
This year I finally did it, and I am so glad that I did! The workshop held Wednesday October 9 in Macon,GA gave myself and my colleagues an excellent excuse to spend more time together. It was a long day but definitely worth the trek. Digitization for Small Institutions was presented from 9 am to 12 noon with a short break in the middle of the session. The two presenters opened by talking about the Digital Library of Georgia (DLG) and right away shared links to resources including a toolkit for Project Managers from the Digital Library Federation (DLF): https://wiki.diglib.org/DLF_Project_Managers_Toolkit. For folks new to using a project management tool, this wiki has an excellent table of recommended software with summaries of each, links to them and pros and cons side by side. Many of the tools you expect to find are here (Jira, Asana, Trello, Slack, Google Suite) although I was personally disappointed that KanbanFlow was not included (insert sad-face emoticon here), there were a few I had not yet heard of or tested out which is ALWAYS exciting.

Photo from my messy notes of a favorite, useful visual.
In the first hour I quickly learned more about DLF, DLG and DPLA (Digital Public Library of America). This was an extremely interesting portion of the workshop that served as the backdrop for the rest of the session's more detailed "how to" segments. Although I had heard of and visited each of the aforementioned DL sites before it had been quite a while since I had taken a moment to just learn more about them and familiarize myself with the "why" of each site and their respective purposes. This seemed particularly relevant after I returned from the conference as we prepared for Open Access Week just a few weeks later. I did not realize how many wonderful resources DLF made available for free online. The project manager toolkit wiki is invaluable, and even if you are not working on projects that will eventually feed up into a DL site, the kit contains so many best practices and tips that it could be useful for many types of digitization projects. One such best practice was this 5-step process (as seen in my messy note photo here): 1. Selection & Planning, 2. Metadata Creation, 3. Prep & Scanning, 4. Post-Processing (crops & edits), 5. Ingest & Preservation (into institutional repository). Before we had a short intermission the attendees were divided into break-out groups of 3 to 4. In this form we discussed why we were there, what projects we were undertaking and what our role was at our institution. Another takeaway takes me back to what I love so much about GLC: there were more staff than librarians in attendance, and a surprising number of public library or museum attendees.

Slide dissecting "Title"
For the rest of the workshop we were shown workflow charts (I LOVE a good visual aid for wrapping my head around a process and grasping a project's big picture) and given what might as well have been a micro-course on metadata terms with a focus on descriptive data, and specifics on Qualified Dublin Core. There was even a little LinkedData talk! What was most helpful about this section were the slides that included specific examples of Title fields. You know a session is worthwhile when you can take that nugget of info back and start using it immediately at work when you return. This was that particular nugget for me!


Hands-on Digitization Station
I was able to share in my breakout group and with the entire group of presenters out loud the challenges of a certain project I have been collaborating on in our library for properly and efficiently archiving thousands of photographs. Lucky for me our project is dealing with media that is already digital, and I already have a space that exists and is ready for hosting the images and metadata (Digital Commons). It was super cool to hear the stories and challenges of others, including what types of media they are digitizing, organizing and archiving to make accessible to their patrons. Not everyone has a repository in place, and not everyone has the staff or tools to achieve their goals right away. This workshop also provided a hands-on station to practice digitization before you left the room. I love that the session enabled everyone, even those interested in the topic (lots of MLS students were there too) but not currently working in a place or role that allows them to get their hands dirty to do just that!

I left the workshop feeling inspired and with an added confidence for the project waiting for me back in the office. Many of the tips I gained from the workshop I am currently utilizing this very week. I had such a wonderful experience that I will certainly sign up for future pre-conference workshops next year! In particular I have enjoyed taking part in GLA's interest groups like Technical Services and Information Technology, and their division sponsored activities like the Academic Library Division, the New Members Round Table Division, and the Paraprofessional DivisionWhat local, state or regional organizations would you recommend to AALL TS-SIS and OB-SIS members who may be from the same area of the States that you are? Share with us in the comments below and link to the association, group or conference!

Thursday, June 27, 2019

Reflections from CALICon19: Two Best Sessions

Looking out over downtown Columbia, SC #CALICon19

Some talk has been floating around in My Communities for covering conferences that may relate to TS, OBS and even CS members. There is understandably a lot of variation in many member job duties and with that plenty of room for overlap of these SIS individuals. Coming from an I.T. department position before my current role this makes total sense, and I can see the benefit to many of us not only having backgrounds in computer science and other technical fields, but also the advantages to continuing education in those areas as library professionals. This is where CALICon comes in.

With Web Developer Leslie Grove at CALICon19
Many of us are familiar at least in some way with CALI the organization (a.k.a. Computer Assisted Legal Instruction). They provide our law students with extremely helpful study aids, plus have resources that help faculty members with all sorts of things. Librarians fit into this section of folks CALI has resources for too. I first heard of CALICon a few years ago when my friend and web developing office-mate Leslie flew to Denver, CO to co-present with our Information Technology Librarian Jason. Their talk titled "Enough to be Dangerous: 00000110 Things Every Beginner Needs to Know about Coding", gave a snapshot of the main programming language fundamentals to the varied CALICon audience. What a neat conference I thought at the time. I had only been to either Drupal Camps full of I.T. guys or state librarian conferences which were full-on librarian attendees, and the idea of a conference that brought librarians, tech-heads and faculty together sounded... well downright phenomenal!

A couple years later I lucked into CALICon coming to Atlanta. Being in Athens, GA it was a short drive. I presented on infographics, and realized I was right - CALICon is pretty amazing. The unique mixture of attendee's makes for interesting discussion and highly useful content that naturally lends itself to collaborative relationships. In true tech-event fashion CALICon live streams all of the sessions and at hyper-speed uploads them all for streaming on YouTube. So, if you have never been to CALICon before, I encourage you to consider it next year. One only has to browse the CALICon playlists of session videos to wonder why everyone doesn't attend.

This brings me to my top 2 sessions from #CALICon19 which I felt would be most useful to  TechScans followers:
  • Leveraging eResources for Affordable Course Materials - Mary and Lisa were excellent presenters who didn't just share something cool (maybe their topic wasn't the flashiest on the schedule) but certainly brought one of the more relevant sessions for me throughout CALICon's two-day whirlwind. What institution isn't interested in saving money for their law students? What library doesn't grapple with ways to make things more cost-effective? This session not only discussed measures that would greatly benefit students but also ideas for faculty members who want to publish their own course content. In this session I learned about lulu.com (CALI actually uses them to publish their books! SUPER affordable, 600+ page books for around $25 shipped!), Powernotes, H20 open casebook platform and more. The presenters even shared strategies for liaising with your registrar office and faculty members to offer alternatives before or alongside booklists, and how they reviewed their own booklists from past semesters to locate and suggest cost-saving measures for specific courses. 
John Presents at CALICon19
  • Automating Processing and Intake in the Institutional Repository with Python - Wow, just wow is all I could say after this session. Most of us deal with our IR in some form or another. As my own role with our Digital Commons site continues to increase, I went into this session with high hopes and seated next to our law school web developer (the office-mate mentioned before), and we were not disappointed. If you have ever manually entered items into your own IR one at a time as I typically do, you smile at the prospect of batch loading. With a large project of archiving old photos in our own IR looming I have been postponing preparing my own spreadsheets - I know it will be tedious and a worm hole of a project. After John's session I am SO glad I waited. My colleague, the coding goddess, and I sat in awe of the automation John was sharing. I was pleasantly rejuvenated leaving the session with a collaborative game plan which I am happy to say we are already making great progress on. Although the presenter's project was with Law Journals and pulling content from PDF's, our own is actually much simpler since we are pulling titles, image URLs and (hopefully) basic descriptions. By far this session left me feeling the most excited about returning to work with something we could instantly put to use.
Click on the session hyperlinked titles for slides and streaming video. Did you attend CALICon too? What were your favorite sessions or biggest takeaways? Find other sessions from CALICon 2019, or past years in CALIorg's YouTube Playlists.

Wednesday, June 19, 2019

BIBFRAME goes International

A recording of the Library of Congress webcast BIBFRAME Goes International. 2019. Video. https://www.loc.gov/item/webcast-8682/ presented April 2, 2019 has been made available. A number of speakers addressed experimentation and implementation of BIBFRAME and/or linked data concepts in Europe, the United States, Asia and Australia/New Zealand.

Some highlights:

Kungliga biblioteket, The Swedish National Library of Sweden has a production BIBFRAME based union catalog available for exploration. They are actively seeking a path out of the MARC environment.

Judith Cannon spoke at length about the PCC/LD4P grant funded group. Seventeen selected PCC libraries are working in a "sand box". Metadata will created and saved using "Sinopia", a linked data platform developed by Stanford University. More information about the project and its goals is available at https://wiki.duraspace.org/display/LD4P2/LD4P2+Project+Background+and+Goals. The Library of Congress is developing initial training material based on LC's BIBFRAME editor. It is not clear when these tools might be available for non-participating libraries to play with.

Paul Frank and Jodi Williamson spoke about Share VDE, a collaboration with Casalini Libri focused on converting MARC bibliographic data to linked data. "VDE" stands for Virtual Discovery Environment. The environment is available for exploration at http://www.share-vde.org/sharevde/clusters?l=en.

Hong Kong University of Science and Technology has an experimental Bibliographic Linked Data Learning Platform. With this tool, you can view and compare bibliographic data presented in different serializations, plus information about the work contextualized using Wikidata knowledge cards. The site also has an experimental SPARQL query form that can be run against their bibliographic data.

The National Library of New Zealand has made their Ngā Upoko Tukutuku / Māori Subject Headings available as linked data as an aid to bibliographic description centered on a  Māori world view.





Friday, March 29, 2019

Library of Congress BIBFRAME progress

A recent ALCTS webinar "Library of Congress BIBFRAME progress" provided information on the current state of BIBFRAME development. Topics included fiscal year 2019 goals and achievements, an exploration of issues mapping MARC to BIBFRAME to MARC, explication of the issue of "blank nodes", and developments in LC's Linked Data Services.

LC is particularly interested in mapping data both into and out of BIBFRAME to eliminate the need for staff participating in the BIBFRAME pilot to do double work. Currently, participating staff are required to describe a resource in BIBFRAME, then re-describe it in MARC. It will be necessary to provide full MARC and BIBFRAME resource descriptions for the foreseeable future. Sally McCallum described several complicated modeling issues that must be resolved before duplicate work can be eliminated. Many of these issues are related to modeling differences; MARC is a "unit record model" with data both from and about the resource integrated into a record, BIBFRAME splits the data about a resource into Works (RDA work/expression), Instance (RDA manifestation), and Items. Decisions are needed on how to present a BIBFRAME work as a MARC work. Use of vernacular scripts and URIs present additional issues. URIs are present in BIBFRAME descriptions in areas that are not currently supported in MARC. Use of URI's to represent concepts at the field level is pretty straight forward, but mapping of headings with qualifiers can be problematic. The goal is to produce structurally sound MARC records from BIBFRAME.

Kevin Ford addressed the issue of "anonymous resources", also known as "blank nodes". He described anonymous resources as a "fact of life" in the context of raw data transformations. Although it would be nice if all data points and concepts had URIs, not everything rises to a level where an entity is willing to mint and maintain a URI. LC is tackling some of this by creating an experimental "providers" file of publishers, available at id.loc.gov/entities/providers.

This presentation is available via the ALCTS YouTube channel at https://youtu.be/YltipGeoJ5Q. ALCTS webinars are generally made available via the ALCTS YouTube channel six months after initial presentation. The Library of Congress makes presentations about BIBFRAME available via their Bibliographic Framework page.


Thursday, February 21, 2019

New Updates to MarcEdit


Creator Terry Reese has been investing time and energy into upgrading the already powerful MarcEdit 7 Editor to fit the needs of users better. By working to improve how manual and global updates are carried out, Rees has improved how fast records load, how changes are tracked (undo!), and rewritten the code to make further edits to the program easier. Read more about the changes that have been made here: https://blog.reeset.net/archives/2762

As of Feb 18, 2019, Reese has also added a custom report writer that can, “search for specific data, either as a match case or regular expression, and return back a report noting # of times in the file and # of records.” For more about the new tool and an example: https://blog.reeset.net/

Wednesday, December 19, 2018

Coming to terms with the IFLA LRM

In early December, ALCTS sponsored a webinar entitled The IFLA LRM Model: an introduction presented by Thomas M. Dousa of The University of Chicago Library. The webinar attempted to distill the concepts embodied in the ILFA Library Reference Model for an audience just beginning exploration of the model. Since the revised RDA Toolkit is organized in alignment with IFLA LRM entities, understanding of the model should aid use of the new toolkit.

Dousa explained that the IFLA LRM represents a harmonization of the three conceptual models sometimes referred to as the "FRAMILY", that is FRBR (Functional Requirements for Bibliographic Records), FRAD (Functional Requirements for Authority Data), and FRSAD (Functional Requirements for Subject Authority Data).In the training period leading up to the introduction of RDA, many of us spent time wrestling with the FRBR WEMI model and the FRBR user tasks of Find, Identify, Select and Obtain.

The LRM presents an expanded suite of user tasks:
  • Find - bring together information about one of more resources of interest ...
  • Identify - clearly understand the nature of the resources found and distinguish between similar resources
  • Select - determine the suitability of the resources found ...
  • Obtain - access the content of the resource
  • Explore - discover resources using the relationships between them, placing the resources in context
And a consolidated list of entities:

  • Res
  • Work
  • Manifestation
  • Expression 
  • Item
  • Agent
  • Person
  • Collective Agent
  • Nomen
  • Place
  • Time-Span
These entities function in an "is-a" hierarchy, where entities inherit the characteristics of entities further up in the hierarchy. Practically, all entities are subcategories of "res", and "person" and "collective agent" are subcategories of "agent". 

The definition of "work" has been adjusted to read "the intellectual or artistic content of a distinct creation". It should be noted that works are modeled as coming into existence with the creation of an initial expression; there is no work without at least one expression of the work. The definition of "expression" has been adjusted to account for simultaneous creation with a work; the definitions of manifestation and item have also been adjusted. 

Some additional things to keep in mind include the definition of "person" in a way that prohibits the treatment of fictional beings as persons , the concept of "nomen" defined as "an association between an entity and a designation that refers to it", and the idea of a "representative expression" as essential to characterizing a work.

The concept of relationships is central to the LRM. Currently 36 have been declared in the format [Entity A]<Relationship>[Entity B]. Relationships among the various WEMI entities form the core of the model.

How any of this will play out in the daily work of bibliographic description remains to be seen. The RDA Steering Committee has yet to finalize revisions to the RDA Toolkit, but it is my understanding that many cataloging policy decisions will be governed by application profiles.

As a reminder ALCTS webinars are made available at no cost on the ALCTS Youtube channel six months after original presentation.

Friday, November 2, 2018

What's up with identity management?

A recent post The coverage of Identity Management work by Karen Smith-Yoshimura in OCLC's Hanging Together blog highlights developments in the probable shift in cataloging practice from "authority control" to "identity management". To put it most simply, our efforts to differentiate creators and correctly correlate their output would shift from constructing a unique text string for each entity to associating the entity with a unique identifier in the form of a URI. Movement towards identity management specifically aligns with the PCC's Strategic Direction 4 "Accelerate the movement toward ubiquitous identifier creation and identity management at the network level"  (https://www.loc.gov/aba/pcc/about/PCC-Strategic-Directions-2018-2021.pdf, page 5).

The Program for Cooperative Cataloging's ISNI Pilot  represents one venue to explore the possibilities of identity management in the context of cataloging. Association of creators with URIs will ease the transition of bibliographic data into a BIBFRAME/linked data environment. The presentations given at the PCC Participant's meeting at ALA Annual in New Orleans provide an overview of the project and examples of project participant's experiences.

Identity management also has the potential to facilitate authority control in the context of journal literature and institutional repositories. How should catalogers provide authority control for journal article authors? Name identifiers in the linked data world (Cataloging & Classification Quarterly 54:8, p. 537-552 (2016) examined the possibilities for using several sources of author identifiers available through international authority databases.  ORCID recently invited feedback on a draft recommendation for ORCID in repositories and is evaluating the use of identifiers for organizations. A recent paper published by JISC explores the potential of Persistent Identifiers to track scholarly work through the research life-cycle, linking the work of researchers with institutions, funding and publication. The focus of the paper is on OA workflows, but the use of PIDs should be applicable across both OA and paid publications.







Friday, September 14, 2018

BIBFRAME Update Forum at the ALA Annual Conference 2018



https://www.loc.gov/bibframe/news/bibframe-update-an2018.html

A BIBFRAME update forum was held at the 2018 ALA Annual Conference with presentations from institutions reporting on projects underway.
Jodi Williamschen, Library of Congress, gave an update on BIBFRAME Pilot 2.0.  She reported that recent infrastructure improvements at LC have been made with the addition of servers and software updates.  The BIBFRAME database, updated daily, contains over 17 million MARC records that have been converted to BIBFRAME Works.
A BIBFRAME 2.0 Implementation Register https://www.loc.gov/bibframe/implementation/register.html is available on the LC website.  Located here is information about a project undertaken at the University of Illinois at Urbana-Champaign Library (UIUC) that focused on creating an interface and converting 7,829 Dublin Core items to BIBFRAME 2.0.  A link is provided to the UIUC Bibframe search interface http://sif.library.illinois.edu/bibframe/search2.php.
A presentation by Tiziana Possemato, Casalini Libri - @Cult, From MARC to BIBFRAME in the SHARE-VDE project, highlighted a collaborative linked data endeavor developed by Casalini Libri (European bibliographic and authority data provider) and @Cult (ILS and Discovery tool provider).  Initial input for the project was received from sixteen North American Research Libraries.
Jeremy Nelson, Metadata & Systems Librarian at Colorado College and co-founder of Knowledgelinks.io presented a model for using BIBFRAME in a multi-institutional projects.  The project known as Plains to Peaks collective attempts to unite isolated digital collections located across Colorado and Wyoming into one platform.
Nathan Putnam, Director, Metadata Quality, OCLC discussed the OCLC Research process in converting approximately 11 million MARC records to BIBFRAME 2.0.  Through the process the team learned the importance of Work IDs and URI.  OCLC remains committed to working with LC to support development of BIBFRAME
For links to individual presentations and further information see the BIBFRAME webpage at the Library of Congress website https://www.loc.gov/bibframe/news/bibframe-update-an2018.html

Monday, July 2, 2018

Stanford Libraries Awarded Grant to Implement LD Environment


The Andrew W. Mellon Foundation has awarded Stanford Libraries a $4 million grant to lead an effort to integrate library data into the greater Web via linked data. Stanford will be partnering with Cornell, Harvard, and the University of Iowa to implement a prototype environment and tools over the next two years. A deliberate partnership with the Program for Cooperative Cataloging (PCC) and the Library of Congress has been included in the project, allowing for an expansion of the number of libraries that will be able to implement linked data.

More details can be found in the press release on Library Technology Guides at https://librarytechnology.org/pr/23584.