论文信息 - Separating the wheat from the chaff: Identifying key elements in the NLA .au domain harvest

Separating the wheat from the chaff: Identifying key elements in the NLA .au domain harvest

In 2005 and 2006 the National Library of Australia (NLA) carried out two whole-domain web harvests which complement the selective web archiving approach taken by PANDORA. Web harvests of this size pose significant challenges to their use. Despite these challenges, such harvests present fascinating research opportunities. The NLA has provided Charles Sturt University's POA (Preservation for Ongoing Accessibility) research group with access to these web harvests and associated keyword indexes. This paper describes the 2006 harvest and uses the example of blogs to address how to identify material within the harvest and determine issues that need further investigation.

Annemaree Lloyd | Ross Harvey | Bob Pymm | Jake Wallis | Geoff Fellows

[1] Elna Saxton. Archiving Websites: A Practical Guide for Information Management Professionals , 2007 .

[2] Shauna-Lee Konrad. Social software in libraries: Building collaboration, communication and community online. , 2008 .

[3] Annemaree Lloyd,et al. Dealing with Digital Collections: Interviews with the National Library and Selected State Libraries of Australia , 2007 .

[4] Julien Masanès,et al. Web Archiving Methods and Approaches: A Comparative Study , 2006, Libr. Trends.

[5] Margaret E. Phillips,et al. What Should We Preserve? The Question for Heritage Libraries in a Digital World , 2006, Libr. Trends.